Wildroot

Queries may use an external AI service. Details

Wildroot Browser

TodayOur model

How good is our model?

An honest report card. Our own model picks the topic of each post, what its author is doing and how worth reading it is. Here is how often it gets that right, compared with Jev, a commercial decision model (TypeSafe System One), asked the same questions with our categories. “Right” means matching our own labelling standard.

topic right
94.5%on the half of Bluesky posts it is most sure about
intent right
92.0%same half; Jev gets 88.5%
cheaper than Jev
~60×$0.0013 vs $0.081 per 1,000 posts

Sure answers are right answers

The model also says how sure it is. When we keep only the posts it is most sure about, it is right far more often. That is why only confident picks are ranked, and everything else goes under “More in this topic”.

Accuracy against the share of posts kept, on 400 Bluesky test posts. Keeping the most confident half: topic 94.5%, intent 92.0% (Jev 88.5%). Keeping everything: topic 71.5%, intent 74.8% (Jev 68.5%).60%70%80%90%100%30%40%50%60%70%100%posts kept, most confident firstIntent (Jev): 88.3% right when keeping the most confident 30%Intent (Jev): 87.8% right when keeping the most confident 40%Intent (Jev): 88.5% right when keeping the most confident 50%Intent (Jev): 84.1% right when keeping the most confident 60%Intent (Jev): 80.8% right when keeping the most confident 70%Intent (Jev): 68.5% right when keeping the most confident 100%Topic (ours): 98.3% right when keeping the most confident 30%Topic (ours): 97.5% right when keeping the most confident 40%Topic (ours): 94.5% right when keeping the most confident 50%Topic (ours): 88.3% right when keeping the most confident 60%Topic (ours): 84.3% right when keeping the most confident 70%Topic (ours): 71.5% right when keeping the most confident 100%Intent (ours): 96.7% right when keeping the most confident 30%Intent (ours): 95.6% right when keeping the most confident 40%Intent (ours): 92.0% right when keeping the most confident 50%Intent (ours): 90.0% right when keeping the most confident 60%Intent (ours): 85.7% right when keeping the most confident 70%Intent (ours): 74.8% right when keeping the most confident 100%Intent (ours) 74.8%Topic (ours) 71.5%Intent (Jev) 68.5%Accuracy against the share of posts kept, on 400 Bluesky test posts. Keeping the most confident half: topic 94.5%, intent 92.0% (Jev 88.5%). Keeping everything: topic 71.5%, intent 74.8% (Jev 68.5%).60%70%80%90%100%30%50%70%100%posts kept, most confident firstIntent (Jev): 88.3% right when keeping the most confident 30%Intent (Jev): 87.8% right when keeping the most confident 40%Intent (Jev): 88.5% right when keeping the most confident 50%Intent (Jev): 84.1% right when keeping the most confident 60%Intent (Jev): 80.8% right when keeping the most confident 70%Intent (Jev): 68.5% right when keeping the most confident 100%Topic (ours): 98.3% right when keeping the most confident 30%Topic (ours): 97.5% right when keeping the most confident 40%Topic (ours): 94.5% right when keeping the most confident 50%Topic (ours): 88.3% right when keeping the most confident 60%Topic (ours): 84.3% right when keeping the most confident 70%Topic (ours): 71.5% right when keeping the most confident 100%Intent (ours): 96.7% right when keeping the most confident 30%Intent (ours): 95.6% right when keeping the most confident 40%Intent (ours): 92.0% right when keeping the most confident 50%Intent (ours): 90.0% right when keeping the most confident 60%Intent (ours): 85.7% right when keeping the most confident 70%Intent (ours): 74.8% right when keeping the most confident 100%Intent (ours) 74.8%Topic (ours) 71.5%Intent (Jev) 68.5%
Bluesky test set: 400 posts labelled independently by two Claude AI labellers under our written guideline, with disagreements adjudicated. “Right” means matching that labelling standard. Hover a point to see its value.
Show the numbers
Posts keptTopic (ours)Intent (ours)Intent (Jev)
30%98.3%96.7%88.3%
40%97.5%95.6%87.8%
50%94.5%92.0%88.5%
60%88.3%90.0%84.1%
70%84.3%85.7%80.8%
100%71.5%74.8%68.5%

Against Jev, on Bluesky

Bluesky test, 400 posts
MeasureOursJev
Topic, all postsTie (not significant)71.5%73.3%
Topic, most confident halfTie94.5%94.5%
Intent, all postsOurs +6.3 (significant)74.8%68.5%
Intent, most confident halfOurs +3.5 (not significant)92.0%88.5%
Quality ranking (Spearman)Ours +0.047 (significant)0.7110.664
Junk detection (AUC)Tie0.9470.944

On Bluesky we beat Jev on intent and on ranking quality, tie on topic, and cost about sixty times less to run.

Where we are weaker

  • Mastodon. Jev is ahead on Mastodon topics (3.2 points across all posts on our latest fresh test) and on ranking quality. We are level with it on the posts we are most sure about and on intent. Our junk filter scores higher, mostly because it also looks at the account.
  • The unsure half. Across all posts, topic is right 71.5% of the time and intent 74.8%. Those are the posts we don’t rank.
  • Small tests. 400 Bluesky and 592 Mastodon test posts leave a few points of uncertainty either way, which is why we say when a gap is too small to be sure of.
Mastodon test, 592 posts from the day after training
MeasureOursJev
Topic, all postsJev +3.2 (significant)82.6%85.8%
Topic, most confident halfJev +2.0 (not significant)97.0%99.0%
Intent, all postsOurs +4.2 (not significant)74.5%70.3%
Quality ranking (Spearman)Jev +0.077 (significant)0.5670.644
Junk detection (AUC)Ours ahead (not tested for significance)0.9080.821

Mastodon has its own topic, intent and quality models, retrained with Mastodon examples added to the Bluesky ones, plus a junk and automation filter that also looks at the account. The two topic rows are for the model that ran when this test was scored. The topic model live since 28 September adds a photography fix: on these same posts it gets 83.3% right, and when it says Photography it is right 85% of the time instead of 70%. We used these posts to design that fix, so this is a check, not a fresh result; a fresh test is planned.

How it learned

It learned from 14,068 public Bluesky posts, labelled by Claude models following our written guideline (v9), with hidden check posts in every batch. Its confidence is tuned on 851 more posts it never trained on. More labels still help, but less each time:

Accuracy on the most confident half, by number of training posts
Training postsTopicIntent
1k88.5%86.8%
4k92.8%89.8%
8k94.2%90.8%
14k94.5%92.0%

It runs on one small server. The only paid step turns each post into its meaning fingerprint, the same fingerprint our similar-post engine uses.

Model cand-v1.1. Sources: our experiment log, entries E70, E74, E75, E93 and E98 (25–28 September 2026).