How this is collected
This is an index of what language models are good and bad at, organised by capability rather than by what is new. Every claim links to the sources it came from, and where the evidence disagrees, both sides are kept.
It is incomplete and partly unreviewed. That is worth saying plainly on the way in, rather than leaving you to discover it. The point of the page you are reading is that you can judge any entry for yourself, because how it got here is recorded on it.
Where entries come from
A nightly job queries OpenAlex for new arXiv papers, filters them to work that is plausibly about a capability tracked here, and asks a model to judge whether each paper is genuinely about that capability and whether it supports or cuts against something already filed. Papers that survive get read in full, and a model drafts a proposed claim from them.
Drafts are not endorsed. They enter the catalog marked pending review, and the rule for those is: visible everywhere, authoritative nowhere. They appear in lists, counts and the contested view like anything else, because an index is useful before it is verified and hiding half of it would be a lie by omission. They are excluded only where a claim would silently decide something — whether a technique is judged to work, and the internal scorecard that tracks whether this catalog is getting things right.
Other entries were filed by hand, usually from well-known papers, when the structure of the catalog was being worked out.
Right now: 82 of 167 claims are pending review, and 82 were drafted by a model rather than written by a person. Most may stay that way. Reviewing is slow and there is more worth indexing than one person can check, so “unreviewed” is a normal permanent state here, not a queue waiting to be cleared.
What a claim carries
- Its scope, inside the sentence. Which models, which tasks, which setup. A claim without its scope condition gets misapplied, so the condition is written into the statement rather than appended.
- Whether it is durable or perishable. A mechanism claim explains why something happens and tends to outlive a model generation. An observation describes how some model or era behaves and is expected to go stale.
- How well it is backed — one paper, replicated across independent work, argued from how something works without anyone measuring it, or someone’s own observation. This is a category, not a score, and nothing here is ranked.
- Its sources, and their side. Each source is marked as supporting or contesting. A contested claim keeps both, plus a written guess at why the evidence disagrees.
What sources are, and what they are worth
| Kind | Count | Treated as |
|---|---|---|
| paper | 144 | Published or preprint work. Titles are checked against arXiv. |
| vendor-doc | 3 | Guidance from whoever builds the model. Privileged about its own product, weak on efficacy (the evaluations are not visible), and interested. Never counted as replication. |
| post | 8 | Informal writing. Good for what practitioners believe and for negative results nobody publishes. Can raise a claim; cannot strongly back one. |
| observation | 0 | Something noticed directly rather than read. |
Anything that can be edited or deleted after the fact — a post, a vendor page, a PDF served from a repository — is archived here verbatim, and a nightly check re-reads the original and reports when it stops matching. A citation should not be able to quietly stop meaning what it meant.
What this gets wrong
- Coverage is uneven. Some capabilities have several claims and some have none. A thin page means nobody has filed here yet, never that there is nothing to know.
- Machine-drafted claims can be wrong in confident-sounding ways. Numbers in a draft are checked against the source text automatically, and anything that fails is flagged on the claim — but a fluent, plausible, mistaken summary passes that check.
- Almost nothing here has been contested by anyone outside. Both sides of a contested claim were assembled by the same reader. That is the weakest part of the whole thing.
- The index is one person’s reading. What gets filed reflects what one person went looking for, which is a bias no amount of structure removes.
None of that is a reason to present the contents with more confidence than they have. It is the reason every entry shows its sources, its backing, and how it arrived — so you can decide, rather than being asked to trust.
Everything is plain YAML in a public git repository, so the history of any entry — when it was filed, what changed, and why — is readable. See the repository.