Slide creation benchmark
The frontier of AI slide creation
Six frontier models completed the same 150 slide creation tasks. We compared output quality, speed, cost, token use, and performance as task complexity increased.
- Frontier models
- 6
- Shared successful tasks
- 139
- Selected AI ratings
- 6,255
TL;DR: GPT-6 Astra is by far the best slide creation model in the Slidely harness. It beats Fable 5.1 while costing less than half as much. It is a quantum leap over GPT-5.6 Sol.
Methodology in brief
Detailed methodology is available below.
We gave every model the same task: create a slide from a layout, content, visual reference, and ASCII plan, while following the layout even when some adaptation was required. The dataset contains 150 samples split across medium, high, and very-high-density slides. Quality results use the 139 tasks completed successfully by every model.
Judging slide creation proved to be extremely difficult. It is not only about visual appeal. A useful slide must follow the layout, group shapes correctly, keep text readable without overlaps, use space well, and create a clear visual hierarchy.
We tried several approaches and compared them with expert human choices on a smaller sample. Individual rubric scores were inconsistent, even with a detailed rubric. Judges that saw many outputs at once often missed important details. Agentic judges that inspected every input helped, but were expensive and still sensitive to the scoring scale.
We settled on blind pairwise comparisons with ties allowed. Three AI reviewers, Gemini 3.7 Flash, GPT-5.6 Sol, and Fable 5.1, compared two anonymous slides at a time. We pooled their decisions and fit a Bradley-Terry model to get a relative quality score. Gemini 3.7 Flash is fixed at 1000 as the reference point. The choice of anchor does not change the ranking or the distance between models.
The results correlated well with expert human choices on the smaller sample. Every creation model used the high reasoning setting, and caching was enabled for all models.
Results and analysis
Overall pooled Arena score
Bradley–Terry scores pooled across the three selected AI reviewers, with Gemini 3.7 Flash anchored at 1000.
GPT-6 Astra performs best in terms of quality, with an Arena score of 1090.5, beating Fable 5.1 by 40 points. Until this generation, OpenAI models performed relatively poorly compared with models from Google and Anthropic in our benchmark and in practical use. Astra reverses that pattern and shows extraordinary spatial reasoning.
Qualitatively, Astra is very good at adapting content without losing the layout's underlying structure. It uses larger and more consistent font sizes when space is available, which has been a persistent problem with every other model we tested. It also handles text contrast and purposeful icon use better. The results feel smart and, more importantly, usable.
As slide density rises, the difference between models becomes more pronounced. Astra and Fable win more often than the Gemini models on higher-density slides.
Pooled Arena score by slide density
Quality scores for medium, high, and very-high-density tasks, independently anchored at Gemini 3.7 = 1000.
Astra leads all three density groups and reaches its highest score, 1113.7, on very-high-density slides. Opus and Fable also improve relative to Gemini 3.7 as slides get harder. Gemini 3.8 moves in the opposite direction, falling from 1013.0 on medium slides to 985.8 on very-high-density slides.
Cost per slide creation task
Median cost per successful slide creation task
Internal billed credits converted to equivalent USD costs for the same 139 shared cases.
Gemini 3.7 Flash is still the cheapest model at $0.24 per median task. Astra costs $1.16, but it is less than half the cost of Fable 5.1 at $2.38 and cheaper than Opus 5 at $1.52. This gives Astra the best combination of quality and cost among the top three models.
Cost increases with density for every model. Astra rises from $1.04 on medium slides to $1.33 on very-high-density slides, while Fable rises from $2.01 to $2.79. The scatter view makes the tradeoff clear: Gemini 3.7 is the economy choice, while Astra buys the best quality without reaching Anthropic's cost.
Time per slide creation task
Median time per successful slide creation task
End-to-end creation time for the same 139 shared cases.
Gemini 3.7 Flash remains the fastest model at 148 seconds per median task. Sol takes 170 seconds and Astra takes 175 seconds. Fable needs 222 seconds and Opus needs 293 seconds.
Astra's time is especially interesting because it leads on quality while finishing close to Sol in time. Density adds only 38 seconds to Astra's median, from 157 seconds on medium slides to 195 seconds on very-high-density slides. Gemini 3.7, Opus, and Fable slow much more on the hardest tasks.
Output tokens per slide creation task
Reasoning and visible output token mix
Median output tokens split into reported reasoning tokens and the remaining generated tokens.
3.7 Flash
3.8 Flash
5.6 Sol
Opus 5
Fable 5.1
6 Astra
Astra uses by far the fewest output tokens. Its median task produces 4,688 tokens, compared with 9,028 for Sol, 14,763 for Fable, 20,719 for Opus, 25,809 for Gemini 3.7, and 47,953 for Gemini 3.8. More output clearly does not mean a better slide.
Astra is also the first model in this test whose reported reasoning tokens are lower than its visible output tokens at the high reasoning setting. There has been speculation around terms such as neuralese and recurrent-depth transformers, but OpenAI's public materials about Astra do not confirm either architecture. The measurable result is simpler: Astra reports much less internal reasoning text while producing the strongest slides.
Other interesting results from our testing
Before finalizing the judging approach, we also benchmarked earlier generations of Google and OpenAI models. We did not test their quality with the current pairwise method, so the historical comparison focuses only on cost, total creation time, and output tokens.
How model economics changed over time
Cost, end-to-end creation time, and output-token use for earlier Google and OpenAI models, shown separately in launch order.
Google: Median cost per task
Google's Flash models repeatedly brought cost and time down after the heavier Gemini 3.1 Pro run. Gemini 3.7 Flash returned to a $0.24 median cost and 148-second median time, while producing roughly twice as many output tokens as Gemini 2.5 Pro.
OpenAI followed a different curve. Time rose from 150 seconds with o3 to 432 seconds with GPT-5.2 and remained high with GPT-5.4. Sol then brought the median back to 171 seconds and cut output to 9,264 tokens, although its median cost rose to $0.71.
Model progress is not one smooth efficiency curve. New releases often improve one operating metric while giving up ground on another. Still, the amount of usable work per dollar has increased substantially as quality has improved. We cannot put one exact number on that historical gain because the older outputs were evaluated using a different quality rubric.
Methodology details
Every model received the same task. For each case, it got the same slide layout, content, ASCII plan, and visual reference. The surrounding system, tools, cached inputs, and creation workflow stayed fixed. Only the model changed.
Scoring the outputs took several attempts to get right. We first asked individual AI judges to score each slide against a rubric covering content, layout, readability, and visual appeal. The scores often bunched near the top, and different judges rarely agreed on exact numbers.
We then tried agentic judges that inspected the slide, layout, and supporting material step by step. We also asked judges to rank several outputs at once. Both approaches helped, but they remained sensitive to the choice of judge and the scoring scale. Showing many slides together also made it easier to miss small but important visual problems.
Blind pairwise comparison worked best. A reviewer sees the reference layout and two anonymous slide images, then chooses the better one or calls a tie. One concrete choice at a time was easier to interpret than an absolute score.
Six models create 15 possible pairs per case. Across 139 cases and three reviewers, that produces 6,255 blind ratings. We combine those choices with a Bradley-Terry model. Each model has an unknown strength, and the calculation finds the strengths that best explain every observed win, loss, and tie. A tie counts as half a win for each model. We then convert the result to a chess-style scale and set Gemini 3.7 Flash to 1000.
The reviewers do not always agree, and we show that instead of hiding it.
Which creation models each judge preferred
Normalized positional scores grouped by judge, with the same model colors used across the article.
Sol and Fable both rank Astra first. Gemini ranks Fable first and Astra fifth. Sol and Gemini place Sol last, while Fable scores Sol even lower. The model order in the middle depends on the reviewer, which is exactly why pooling matters.
Agreement between the selected AI judges
Exact agreement on shared pairwise decisions. Fable made all 12 tie decisions.
Gemini and Sol
Gemini and Fable
Sol and Fable
Gemini and Sol agree on 60.2 percent of decisions. Gemini and Fable agree on 61.3 percent, while Sol and Fable agree on 63.9 percent. Agreement around 60 percent is meaningful, but not strong enough to trust one judge as the final answer.
Do judges prefer models from their own lab?
Preference for each judge's model family in cross-family battles, compared with how the other two judges rated that same family.
Every judge preferred outputs from its own model family more often than the other two judges did. The Gemini judge preferred Google models in 51.6 percent of cross-family matches, compared with 45.7 percent for the other judges. Sol preferred OpenAI models 50.4 percent of the time, compared with 44.0 percent. Fable showed the largest difference, preferring Anthropic models 63.4 percent of the time, compared with 52.6 percent.
This is evidence of a possible family preference, not proof of bias. Each family contains models with different underlying quality. Pooling judges from three labs reduces the chance that one judge's family preference determines the final ranking.
The approach also scales well. When a new model arrives, we do not need to repeat every old comparison. We run the new model on the same cases, compare it with the existing models, and refit the Bradley-Terry scores with the additional results.
We also plan to make the blind arena public so that human reviewers can participate. That will give us a stronger independent signal and make it possible to compare AI and human preferences at a much larger scale.
Appendix
Cache reuse between turns
Cache hit rate after the first turn
Share of input tokens served from cache on repeat model calls across the 139 shared tasks.
The chart measures cached input tokens as a share of all input tokens after the first model call in each task. Fable has the highest repeat-turn cache hit rate at 88.5 percent, followed by Sol at 88.0 percent, Opus at 86.9 percent, Astra at 84.9 percent, Gemini 3.8 at 71.4 percent, and Gemini 3.7 at 60.4 percent.
The long tail behind the medians
The main article uses medians because a few unusually difficult tasks can distort the mean. The p99 value is the second-highest observation in this 139-case sample. The p100 value is the maximum.
| Model | Median time | Mean time | p99 time | p100 time |
|---|---|---|---|---|
| Gemini 3.7 Flash | 148s | 171s | 458s | 471s |
| Gemini 3.8 Flash | 207s | 217s | 384s | 414s |
| GPT-5.6 Sol | 170s | 178s | 353s | 416s |
| Opus 5 | 293s | 306s | 685s | 826s |
| Fable 5.1 | 222s | 241s | 659s | 850s |
| GPT-6 Astra | 175s | 181s | 286s | 290s |
| Model | Median cost | Mean cost | p99 cost | p100 cost |
|---|---|---|---|---|
| Gemini 3.7 Flash | $0.236 | $0.250 | $0.545 | $0.636 |
| Gemini 3.8 Flash | $0.309 | $0.329 | $0.609 | $0.655 |
| GPT-5.6 Sol | $0.712 | $0.736 | $1.212 | $1.227 |
| Opus 5 | $1.518 | $1.574 | $3.200 | $3.664 |
| Fable 5.1 | $2.382 | $2.478 | $5.645 | $6.382 |
| GPT-6 Astra | $1.164 | $1.199 | $1.818 | $1.873 |
| Model | Median output tokens | Mean | p99 | p100 |
|---|---|---|---|---|
| Gemini 3.7 Flash | 25,809 | 26,710 | 48,559 | 60,782 |
| Gemini 3.8 Flash | 47,953 | 48,255 | 68,527 | 68,622 |
| GPT-5.6 Sol | 9,028 | 9,539 | 18,080 | 19,195 |
| Opus 5 | 20,719 | 21,698 | 46,550 | 57,584 |
| Fable 5.1 | 14,763 | 16,309 | 47,151 | 57,947 |
| GPT-6 Astra | 4,688 | 4,933 | 8,419 | 8,959 |
Astra has the tightest time and output-token tail in the group. Fable's maximum task took 850 seconds and cost $6.38, while Astra's maximum took 290 seconds and cost $1.87.
Renders produced
Mean renders produced per final slide
Each render is one visible attempt produced before the final slide was selected.
All models averaged roughly four visible render attempts per final slide. Gemini 3.8 had the highest average at 4.2, while Astra and Sol had the lowest at 3.7. Most p99 values sit between six and eight renders, while Fable had one task that reached eleven.
Pairwise match results
Pairwise wins across all three judges
Total wins across 2,085 pairwise ratings per model on the 139 shared tasks.
| Model | Wins | Losses | Ties |
|---|---|---|---|
| GPT-6 Astra | 1,311 | 774 | 0 |
| Fable 5.1 | 1,172 | 910 | 3 |
| Opus 5 | 1,119 | 964 | 2 |
| Gemini 3.8 Flash | 1,004 | 1,074 | 7 |
| Gemini 3.7 Flash | 994 | 1,082 | 9 |
| GPT-5.6 Sol | 643 | 1,439 | 3 |
Each model has 2,085 pooled outcomes because it appears in five matchups per case across 139 cases and three reviewers.
We need even better benchmarks, and Slidely is working on it
This benchmark is already difficult, but Astra's lead shows why the next version must be harder. Arena scores are relative to this task and this slide creation system. They should not be compared directly with unrelated leaderboards.
Quality is measured only on the 139 cases where every model succeeded. Each model produced one final output per case, so we do not yet measure variation across repeated runs. The density groups contain different cases, which means density trends are observed patterns rather than a controlled experiment. The three judges are also AI models with imperfect agreement, and all three model families appear as both contestants and reviewers.
We are working on harder slide creation cases, more human reviewers, stronger independent judges, and repeated generations that measure consistency. We are also building longer presentation creation tasks that test whether an agent can plan and maintain quality across many slides, much longer runtimes, and higher costs.
Slide creation is only half of presentation work. Our upcoming edit-hard benchmark focuses on difficult slide editing tasks where a model must change an existing slide without breaking the layout, content, or surrounding design. We will publish those results soon.