SlidelyBench
How we compare AI slide-editing quality in Slidely
AI presentation tools are getting much better at producing slides that look polished. But looking polished is not the same as completing a slide-editing task.
Early benchmark results
Score vs. total time
Compare slide-editing quality against the total time required.
User-outcome comparison
Overall win rate by tool
Share of the supplied head-to-head results won by each evaluated tool.
Task
- Convert to Title and Content layout. Remove all projects that have X in FY25, add Halyk and Bank of South Sudan at the bottom, and fix font consistency.
| Required change | |||||
|---|---|---|---|---|---|
| Explicit tasks | |||||
| Remove rows | |||||
| Add rows | |||||
| Change layout | |||||
| Implicit tasks | |||||
| Retain row markers | |||||
| Follow design | |||||
| Overall result | Usable | Unusable | Unusable | Unusable | Unusable |
| Tool output | |||||
A tool can generate a visually attractive result while changing the user’s data, deleting useful assets, ignoring the supplied PowerPoint layout, flattening an editable chart into an image, or redesigning parts of the slide that were never meant to change. The output may look good in a screenshot, yet require more work to repair than the original task would have taken manually.
At Slidely, we are building an AI agent that creates and improves complex, fully editable PowerPoint presentations. It needs to work with native layouts, placeholders, tables, charts, templates, and existing user content, not just render attractive slide images.
That requires an evaluation system designed around the way professionals actually edit presentations.
Why conventional slide evaluation falls short
Most slide evaluations focus on some combination of content quality, visual design, readability, and instruction following. These dimensions matter, but they can produce misleading results for editing tasks.
What the screenshot suggests
Global Performance Overview
Solutions
Q1 2024 Revenue by Region (USD millions)
- Title and Content layout
- Company logos added
- Country flags added
- Table converted to chart
What the PowerPoint file reveals
Global Performance Overview
Q1 2024 Revenue by Region (USD millions)
Fake layout made with
floating shapes
Solutions
Logos recreated with
drawing objects
Flags drawn as
shapes / emoji
Chart pasted as
a flat image
Consider a simple request:
- Apply the Title and Content layout.
- Add company logos.
- Add country flags.
- Convert a table into a chart.
An output might appear to satisfy every instruction. But when we inspect the PowerPoint file, we may find that the layout was visually imitated using free-floating shapes, the logos were manually reconstructed using PowerPoint objects, the flags were drawn using shapes or emoji, and the chart was inserted as a flattened image.
The slide may look convincing in a screenshot. For the user, however, only the first task was completed in a usable form. The remaining elements need to be deleted and recreated. A generic visual-quality score can hide this failure.
Recent PowerPoint benchmarks have made meaningful progress beyond purely aesthetic evaluation. PPT-Eval uses task-specific rubrics, partial credit, penalties for unnecessary changes, and natural-language feedback. PPTArena focuses on in-place edits to real PowerPoint files and evaluates both visual outputs and structural differences. PresentBench shows the value of fine-grained, instance-specific evaluation criteria over broad holistic judgments.
SlidelyBench builds on this direction, but optimizes for a narrower objective: evaluating whether professional slide edits are genuinely usable.
Building SlidelyBench around real editing work
The tasks resemble the instructions people give presentation specialists and AI agents in practice:
“Convert this slide to the new template.”
“Add the company logos and country flags.”
“Turn this table into a chart.”
“Improve the visual hierarchy without changing the content.”
“Remove projects marked X and add these two rows.”
“Fix font consistency.”
“Redesign this slide, but preserve all the existing assets.”
These requests are often short and underspecified. There may be many good visual solutions, but there are still clear boundaries around what the tool was asked to change and what it should preserve.
The benchmark therefore does not compare outputs against one canonical screenshot. Instead, it constructs a task-specific rubric from the original slide, the instruction, and the supplied design system.
The rubric is locked before any outputs are scored. This prevents the evaluator from retrofitting criteria around whichever output looks best.
Turn every instruction into explicit and implicit tasks
Slide-editing instructions are usually short, but the requirements for a usable result are not.
A user may explicitly ask to change a layout, add an asset, or convert an object. They rarely specify every implementation detail: use the native PowerPoint layout, keep the object editable, preserve unaffected content, retain existing assets, and follow the supplied template.
A benchmark that evaluates only the literal wording of the instruction can therefore reward outputs that appear correct but are unusable.
SlidelyBench converts each request into:
- Explicit tasks: The changes directly requested by the user.
- Implicit tasks: The structural, preservation, and usability requirements necessary to complete those changes correctly.
Together, these form a task-specific rubric that reflects what the user actually needs, not just what the instruction literally says.
Prompt
Convert to Title and Content layout. Remove all projects that have X in FY25, add Halyk and Bank of South Sudan at the bottom, and fix font consistency.
Explicit tasks
- 1Apply the Title and Content layout
- 2Remove rows marked X in FY25
- 3Add Halyk and Bank of South Sudan
- 4Fix font consistency
Implicit tasks
- 1Preserve unaffected rows and slide meaning
- 2Retain the user’s icons and legend
- 3Recalculate the total after row changes
- 4Match the supplied template and formatting
Score explicit tasks on usability and quality
Each explicit task is evaluated on two separate questions:
- Was the requested result delivered in a usable form?
- How well was it executed?
A result can look polished without being usable. For example, a table recreated with individual lines and text boxes may resemble a real table, but the user would need to delete it and rebuild it as a native PowerPoint table.
The reverse is also possible. A tool may create the correct editable object but execute the task imperfectly—for example, by omitting a requested row, using an incorrect label, or applying weak formatting.
SlidelyBench therefore assigns each explicit task:
- Task completion: whether the requested result is fundamentally correct and usable.
- Task quality: how accurately and professionally the completed result was executed.
A task is complete when the user can keep the delivered result and correct any remaining issues in place. It is incomplete when the result is missing, materially wrong, or must be replaced or recreated.
Each task contributes:
Task weight × task completion × task quality
This prevents visually convincing but unusable outputs from receiving credit, while still distinguishing between usable results of different quality.
Treat preservation of the user’s work as part of correctness
Completing the explicit tasks is not enough if the tool damages parts of the slide the user did not ask it to change.
SlidelyBench therefore evaluates the implicit requirements separately. It checks whether the output preserves:
- Existing content and data.
- Images, logos, icons, and other user assets.
- Legends, markers, sources, and explanatory context.
- The slide’s meaning and information structure.
- Native PowerPoint objects and editability.
- The supplied template and design language.
Failures are treated as unintended changes and penalized according to how much work the user must do to restore the slide:
- Minor: a small, localized repair.
- Medium: noticeable repair to one section.
- Major: important content, assets, or structure must be restored.
- Unusable: a major section, or the entire slide, effectively needs to be reconstructed.
Related changes caused by the same editing decision are grouped into a single incident rather than counted separately.
The final score is:
Final score = explicit-task score − unintended-change penalties
This ensures that a tool is rewarded not only for making the requested changes, but also for preserving everything the user still needs.
We evaluate the PowerPoint, not just the screenshot
Many properties that matter to professionals are invisible in a rendered image. Two slides can look identical while being fundamentally different PowerPoint files.
SlidelyBench checks whether:
- The supplied slide master is actually used.
- The correct native layout is applied.
- The title is inside the title placeholder.
- Source and subtitle placeholders are used when relevant.
- Optional placeholders are allowed to remain empty.
- Charts are native, editable PowerPoint charts.
- Tables are native PowerPoint tables.
- Logos and icons are real image or vector assets.
- Theme fonts and colors are used instead of hard-coded approximations.
- Master elements have not been recreated using slide-level shapes.
- Objects remain sensibly grouped and editable.
We combine slide renders with PowerPoint object data rather than inferring document structure from appearance.
This reflects the product we are trying to build. Slidely works inside PowerPoint and emphasizes fully editable output, template adherence, autolayout, and spreadsheet-backed charts and tables.
Editability is not an export option added after generation. It is part of whether the task was completed correctly.
Task-specific grading creates stronger separation
A fixed visual-quality rubric asks the same broad questions for every slide. SlidelyBench changes the weighting based on the task.
For a request to add logos, flags, and a chart, the benchmark emphasizes:
- Whether every requested asset was added.
- Whether the assets are authentic and mapped correctly.
- Whether the chart is native and editable.
- Whether the source data was preserved.
For a broad redesign request, it gives more weight to information hierarchy, visual composition, readability, design-language fidelity, and improvement over the original.
For a small formatting request, it heavily penalizes unnecessary structural changes.
This makes the benchmark sensitive to the capability being tested instead of collapsing every task into a generic measure of ‘good slide design.’
In our early comparisons, this creates substantially more separation than holistic visual scoring.
Outputs that look polished but are structurally unusable or destructive fall sharply. Outputs that preserve the user’s work and make precise, PowerPoint-native edits rise.
Aligning offline evaluation with user outcomes
An offline benchmark is useful only if it predicts what users experience in the product.
Cursor describes a similar challenge in coding: an agent output can appear correct to an offline grader but feel worse to a developer using it. CursorBench combines production-derived offline tasks with controlled online evaluation to keep its definition of quality aligned with real work.
We are applying the same principle to presentations. Alongside SlidelyBench, we evaluate production signals such as:
- Whether users accept or discard generated slides.
- Whether they undo or regenerate an edit.
- How much manual repair follows an agent action.
- Whether generated objects remain in the final deck.
- Follow-up instructions required to correct the result.
- User preference between alternative outputs.
The strongest benchmark is not the one with the most elaborate rubric. It is the one whose rankings match the outputs professionals actually choose to keep.
What comes next
The current version of SlidelyBench focuses on [SINGLE-SLIDE / EDITING TASK DESCRIPTION]. We are expanding it toward harder workflows:
- Compound instructions spanning several slide elements.
- Edits across multiple slides.
- Deck-wide consistency checks.
- Template conversion across heterogeneous source decks.
- Linked Excel tables and charts.
- Multi-turn editing sessions.
- Long-running agent workflows.
- Consistency with organization-specific design memories.
We also plan to measure latency and compute alongside quality. The best slide agent is not simply the one that eventually produces the highest-scoring file. It must do so quickly enough to remain useful inside an interactive PowerPoint workflow.
As slide agents become more capable, visually plausible output will become easier to produce.
The harder problem is reliable editing: making exactly the requested changes, preserving everything the user still needs, and leaving behind a PowerPoint file that a professional can continue working with.
That is what SlidelyBench is designed to measure.