How CallTally scores a call
Every number on a record comes from the rules on this page. They are the same for every account, nothing is edited by hand, and any scored call can be disputed.
Formulas at a glance
Horizons are 1 week, 1 month and 3 months after the post.
- Return
- (close at window end − close at entry) ÷ close at entry
- Excess return, long call
- stock return − SPY return
- Excess return, short call
- SPY return − stock return
- Hit
- excess return above zero. Anything else is a miss.
- Hit rate
- hits ÷ scored calls, per horizon
- Excess return for an account
- average excess return of its scored calls, per horizon
- 90% interval
- Wilson score interval around the hit rate, z = 1.645
What counts as a call
A post counts as a call only when all three are true.
- It names a ticker
- With or without a cashtag: $NVDA and NVDA both count. A sector, an index theme or “tech” does not.
- It states a direction
- Long or short, in words a reader can't take two ways: buying, adding, long, selling, shorting, short.
- It is the account's own post
- Written and posted by the account. The post's timestamp is where the call starts.
If the post states a horizon, such as “for the next 3 months”, the record notes it. Every call is still scored at all three horizons, so no account gets to pick the window that flatters it.
Calls are pulled out of posts by an automated extraction step with a strict format: ticker, direction, stated horizon if any, the post quoted word for word, and its timestamp. It runs offline on the operator's machine before anything is published, never in your browser.
What doesn't count
- Vague sector talk
- Views on a sector, the market or the economy without a named ticker.
- Questions
- Asking whether to buy or sell is not saying to.
- Retweets of other accounts
- Retweets are excluded when the posts are pulled. A call only counts in the account's own words.
- After-the-fact posts
- “I told you so” posts about a move that already happened. They point back in time, so there is nothing left to measure.
- Ambiguous posts
- If the ticker or the direction could be read two ways, the post is left out. Precision comes before coverage.
Five made-up posts, and how they are read
Long $NVDA into earnings.
Counts: ticker and direction are both stated.
Shorting TSLA here, this bounce fades.
Counts: ticker and direction are both stated.
Semis look strong this year.
Doesn't count: no ticker.
Is $AAPL a buy at these levels?
Doesn't count: a question is not a call.
Called $AMD at 120. Told you.
Doesn't count: posted after the move.
How one call is scored
Each call is measured three times: 1 week, 1 month and 3 months after it was posted. The benchmark is SPY, the S&P 500 fund, over the same window. That is what a follower could have earned by doing nothing.
The window starts at the first daily close at or after the post's timestamp. It ends at the daily close 1 week, 1 month or 3 months later; when that day has no trading, the last close before it is used. Returns use closing prices and leave out dividends.
Return = (close at window end − close at entry) ÷ close at entry
Excess return, long call = stock return − SPY return
Excess return, short call = SPY return − stock return
A call is a hit when its excess return is above zero, meaning the stock beat SPY in the direction the account called. Anything else, including exactly zero, is a miss. The decision uses the unrounded figure, and records show returns rounded to one decimal.
| Stock close at entry | $50.00 |
|---|---|
| Stock close 1 month later | $53.00 |
| Stock return | +6.0% |
| SPY close at entry | $500.00 |
| SPY close 1 month later | $510.00 |
| SPY return | +2.0% |
Excess return vs SPY, long call: stock return minus SPY return
+4.0%
Result Hit
| Call | Stock | SPY | In the call's direction | Result |
|---|---|---|---|---|
| Long | +6.0% | +2.0% | +4.0% | Hit |
| Long | +1.0% | +2.5% | −1.5% | Miss |
| Short | −3.0% | +1.0% | +4.0% | Hit |
How an account's record is computed
An account's record has one line per horizon. Each line has four figures.
- Hit rate
- Hits divided by scored calls at that horizon. Wherever one hit rate stands for an account, it is the 1-month figure and it says so.
- Excess return vs SPY
- The average excess return of the account's scored calls at that horizon, every call weighted the same.
- Scored calls
- Calls whose window has closed. A call posted two weeks ago counts at 1 week but not yet at 1 month, so the three counts can differ. Its open windows show as “Window still open”.
- 90% interval
- The range the account's true hit rate likely sits in, given how many calls there are. The next section shows how it is computed.
Records carry numbers and quotes only. There are no rankings, no comparisons between accounts and no words about an account beyond its handle and name.
The 90% interval and sample size
A hit rate from 11 calls says much less than one from 176. The interval shows that directly: the fewer the calls, the wider the range. It is a Wilson score interval, which stays inside 0–100% and behaves well when there are only a few calls.
With n scored calls, p = hits ÷ n, and z = 1.645 for 90%:
centre = (p + z² ÷ 2n) ÷ (1 + z² ÷ n)
half-width = z × √p(1 − p) ÷ n + z² ÷ 4n² ÷ (1 + z² ÷ n)
interval = centre − half-width to centre + half-width
For 7 hits in 11 calls: p = 0.636, centre = 0.609, half-width = 0.215, so the interval is 39–82%. These are hypothetical numbers.
| Scored calls | Hits | Hit rate | 90% interval |
|---|---|---|---|
| 11 | 7 | 64% | 39–82% |
| 44 | 28 | 64% | 51–74% |
| 176 | 112 | 64% | 58–69% |
Where the data comes from
Posts come from a one-time pull through the official X API: about a year of each account's own posts, for a curated list of about 20 finance accounts, with retweets excluded. There is no scraping and no bought or resold dataset.
Prices come from a file of daily closing prices for every called ticker and for SPY. Each record shows the date of the data run its figures come from.
Records are fixed at that pull. Posts made after it are not tracked yet; real-time tracking is planned as a paid tier and is not built.
Sample data until the first data run
Until the first data run, every scored account on CallTally is fictional. There are five of them, each handle starts with sample_, and their posts were written for the sample. No real person's record is ever made up. Every sample card and record carries this stamp:
How disputes are reviewed
Every scored call has a Dispute button. You don't need to sign in to use it.
- Press Dispute on the call, choose one reason and, if you want to hear the outcome, leave an email.
- The call shows “Dispute under review.” Its numbers and the account's numbers stay as they are.
- A person checks the original post against the reason you chose.
- If the dispute holds, the call is corrected or removed and the account's record is computed again. If it doesn't, the call stays as scored. Either way, you hear back if you left an email.
Reasons you can choose
- Not a call
- The post isn't a ticker-named call with a clear direction.
- Wrong ticker or direction
- The call is real but recorded against the wrong stock or the wrong side.
- Wrong timestamp
- The post time on the record doesn't match the original post.
- Post deleted or edited
- The original post was removed or changed after it was captured.
- Other
- Anything else that makes the scored call wrong.
Disputes and the emails left with them are never shown publicly. The only public trace is the “Dispute under review” note on the call.