AAP-Bench
AAP-Bench measures one thing: whether a language model can produce a complete academic document that compiles, renders, and holds together as a single argument. It scores LaTeX that builds, diagrams that draw, mathematics that typesets, and whether sections accumulate rather than repeat.
It does not measure general capability, factual accuracy, domain expertise, or writing taste. A model can score well here and be wrong about everything it says.
There is no published edition yet.
What follows is a pilot — real generations by real models, shown because hiding them would be worse, but not a result. The runs came from a build that could not identify its own commit, so we cannot say which version of our pipeline produced them. Each model has fewer than the 15 runs an edition requires, and these are rates over small denominators — one failed diagram can move a figure twenty points. Do not cite this table.
What actually happens in real use
Every paper generated with Auto Academic Paper contributes its outcome here — anyone who uses the tool adds to it, on their own key and their own topic. That makes this an observational record, not a head-to-head: the briefs differ, the topics differ, and one model may simply have drawn easier work. Read it as a field failure rate, not a ranking.
Only counts are published — never anyone's paper. No text, no figures, no titles, no briefs. The query behind this table selects a model name and integers and nothing else, and a test asserts on every build that no paper content can reach it.
| Model | Papers | Diagrams failed | Math failed | Sections degraded |
|---|---|---|---|---|
| Claude-Opus-4.8 | 1 | too few papers to report a rate yet | ||
| deepseek-v4-flash | 6(0 measured) | too few papers to report a rate yet | ||
| deepseek-v4-pro | 2(0 measured) | too few papers to report a rate yet | ||
| gemini-flash-latest | 1(0 measured) | too few papers to report a rate yet | ||
A model needs 10 measured papers before rates are shown. Below that a single bad paper moves a figure by tens of points, and the percentage would mislead rather than inform — so those models are listed with their count and no score. Nothing has cleared that bar yet.
The same brief, side by side
Separately, we run every model against an identical published brief — same task, same tier, same judge. Small numbers, but controlled, so this is the half that supports a direct comparison. These papers are ours, which is why excerpts and diagrams from them can be shown below.
| Model | Runs | Diagrams failed | Sections degraded | Cross-ref | Repetition | Connectivity* | Contradictions* |
|---|---|---|---|---|---|---|---|
| GPT-5.4-Mini | 1 | 0% | 0% | 0.000 | 0.010 | 0.750 | 1.000 |
| Gemini-3.1-Flash-Lite | 1 | 0% | 0% | 0.000 | 0.018 | 0.750 | 2.000 |
* Columns marked with an asterisk are LLM-assisted — scored by the edition's judge model. Everything to their left is deterministic: recomputable from the published raw rows by anyone, offline, with no model involved. If you do not accept LLM-as-judge, the deterministic columns alone still give you a complete ranking.
Not measured yet: mathematics. The product validates every math expression with the same engine that renders it, but the harness that produces these runs bypasses that step, so we have no math figure to report. It is absent rather than shown as zero — a blank we could have drawn as a dash and let you read as "nothing went wrong". It returns when these runs go through the full pipeline.
The same task, in their own words
Every model below was given the identical brief — the task spaced-repetition from the published set — and each excerpt is taken from the same position in each paper: the middle section, first 800 characters. The rule is fixed and published so we cannot show you a model's best paragraph and call it representative.
Each diagram below is drawn by the same engine that renders a finished paper, from the model's own TikZ source, unedited. If one looks wrong — labels colliding, arrows crossing, text overflowing a box — that is what the model produced. We are not correcting it, and a diagram that fails to draw at all is shown failing rather than replaced with a picture.
GPT-5.4-Mini
Beyond Fixed Intervals: A Theoretical Framework for Learner-Specific Adaptive Spaced Repetition Schedules
A Formal Model of Adaptive Interval Scaling The central premise of adaptive interval scaling is that review timing should not be governed by a single canonical schedule, but by a learner-specific estimate of retention risk. This premise is consistent with the broader literature on adaptive educational systems, which emphasizes that learners differ in knowledge, affective state, skills, and behavior, and therefore require individualized modeling rather than uniform treatment (ref_7). At the same time, the practical literature on spaced repetition and spaced learning indicates an important limitation of many mainstream tools: in spite of their operational value, they often implement fixed schedules or weakly parameterized heuristics, whereas the underlying notion of spacing is itself…
Its first diagram, rendered:
Gemini-3.1-Flash-Lite
Beyond Fixed Intervals: A Theoretical Framework for Adaptive Spaced Repetition via Dynamic Forgetting Curve Modeling
The challenge of optimizing retention in Spaced Repetition Systems (SRS) lies in the inherent variability of human memory consolidation. While classical models, such as those derived from Ebbinghaus's seminal work, provide a foundational understanding of the forgetting curve, they often fail to account for the idiosyncratic nature of memory decay (ref_1). As noted in recent literature, the Memory Chain Model (MCM) offers a superior fit to retention data compared to traditional power functions by explicitly modeling memory consolidation processes (ref_1). However, current implementations often suffer from the Fixed Interval Fallacy, where system-wide parameters are applied uniformly to a heterogeneous population of learners, ignoring the capacity constraints of working memory (ref_6) and…
Its first diagram, rendered:
This model drew no diagram for that task.
How a result is defined
An edition is a fixed tuple, and every run carries it. Change any part and it is a different experiment — results from different editions are never pooled and never compared.
- • the task set and its version — the exact inputs, identical for every model
- • the writer model — the contestant
- • the judge model — pinned, disclosed, and never a contestant; a model that grades its own paper produces a number no disclosure repairs
- • our own pipeline's commit — so that a fix on our side never reads as a model improving
- • the tier and review depth — pinned identically for everyone, as is reasoning effort
Editions present in the data below:
aap-bench/v1 | minimal | quick | judge=Claude-Sonnet-4.6 | code=UNPINNED
Check us
We sell paper generation and this benchmark runs on the product we sell. That is exactly why the underlying data is published rather than an average we computed. The rows include our own failures. Nothing here has to be taken on trust.
- • The task set — the exact inputs, verbatim
- • Every raw run — the frozen record of each generation, not a summary
- • The method — protocol, sampling, and the judge policy
- • The reporter — deterministic and offline; run it on the raw rows and you get this table