TodayOur model
How good is our model?
An honest report card. Our own model picks the topic of each post, what its author is doing and how worth reading it is. Here is how often it gets that right, compared with Jev, a commercial decision model (TypeSafe System One), asked the same questions with our categories. “Right” means matching our own labelling standard.
- topic right
- 94.5%on the half of Bluesky posts it is most sure about
- intent right
- 92.0%same half; Jev gets 88.5%
- cheaper than Jev
- ~60×$0.0013 vs $0.081 per 1,000 posts
Sure answers are right answers
The model also says how sure it is. When we keep only the posts it is most sure about, it is right far more often. That is why only confident picks are ranked, and everything else goes under “More in this topic”.
Show the numbers
| Posts kept | Topic (ours) | Intent (ours) | Intent (Jev) |
|---|---|---|---|
| 30% | 98.3% | 96.7% | 88.3% |
| 40% | 97.5% | 95.6% | 87.8% |
| 50% | 94.5% | 92.0% | 88.5% |
| 60% | 88.3% | 90.0% | 84.1% |
| 70% | 84.3% | 85.7% | 80.8% |
| 100% | 71.5% | 74.8% | 68.5% |
Against Jev, on Bluesky
| Measure | Ours | Jev |
|---|---|---|
| Topic, all postsTie (not significant) | 71.5% | 73.3% |
| Topic, most confident halfTie | 94.5% | 94.5% |
| Intent, all postsOurs +6.3 (significant) | 74.8% | 68.5% |
| Intent, most confident halfOurs +3.5 (not significant) | 92.0% | 88.5% |
| Quality ranking (Spearman)Ours +0.047 (significant) | 0.711 | 0.664 |
| Junk detection (AUC)Tie | 0.947 | 0.944 |
On Bluesky we beat Jev on intent and on ranking quality, tie on topic, and cost about sixty times less to run.
Where we are weaker
- Mastodon. Jev is ahead on Mastodon topics (3.2 points across all posts on our latest fresh test) and on ranking quality. We are level with it on the posts we are most sure about and on intent. Our junk filter scores higher, mostly because it also looks at the account.
- The unsure half. Across all posts, topic is right 71.5% of the time and intent 74.8%. Those are the posts we don’t rank.
- Small tests. 400 Bluesky and 592 Mastodon test posts leave a few points of uncertainty either way, which is why we say when a gap is too small to be sure of.
| Measure | Ours | Jev |
|---|---|---|
| Topic, all postsJev +3.2 (significant) | 82.6% | 85.8% |
| Topic, most confident halfJev +2.0 (not significant) | 97.0% | 99.0% |
| Intent, all postsOurs +4.2 (not significant) | 74.5% | 70.3% |
| Quality ranking (Spearman)Jev +0.077 (significant) | 0.567 | 0.644 |
| Junk detection (AUC)Ours ahead (not tested for significance) | 0.908 | 0.821 |
Mastodon has its own topic, intent and quality models, retrained with Mastodon examples added to the Bluesky ones, plus a junk and automation filter that also looks at the account. The two topic rows are for the model that ran when this test was scored. The topic model live since 28 September adds a photography fix: on these same posts it gets 83.3% right, and when it says Photography it is right 85% of the time instead of 70%. We used these posts to design that fix, so this is a check, not a fresh result; a fresh test is planned.
How it learned
It learned from 14,068 public Bluesky posts, labelled by Claude models following our written guideline (v9), with hidden check posts in every batch. Its confidence is tuned on 851 more posts it never trained on. More labels still help, but less each time:
| Training posts | Topic | Intent |
|---|---|---|
| 1k | 88.5% | 86.8% |
| 4k | 92.8% | 89.8% |
| 8k | 94.2% | 90.8% |
| 14k | 94.5% | 92.0% |
It runs on one small server. The only paid step turns each post into its meaning fingerprint, the same fingerprint our similar-post engine uses.
Model cand-v1.1. Sources: our experiment log, entries E70, E74, E75, E93 and E98 (25–28 September 2026).