SlidelyBench

How we compare AI slide-editing quality in Slidely

AI presentation tools are getting much better at producing slides that look polished. But looking polished is not the same as completing a slide-editing task.

Early benchmark results

Score vs. total time

Compare slide-editing quality against the total time required.

ScoreTotal timeClaude takes 5× longerClaude Design scores10× lower88.3%11m 41s66.8%62m 03s54.8%59m 13s20.7%39m 26s8.7%11m 50sSlidely, BalancedSlidely, Balanced: Score 88.3%, Win rate 75%, Total time 11m 41sClaude, FableClaude, Fable: Score 66.8%, Win rate 20%, Total time 62m 03sCopilot, Opus 4.8Copilot, Opus 4.8: Score 54.8%, Win rate 5%, Total time 59m 13sChatGPT, HeavyChatGPT, Heavy: Score 20.7%, Win rate 0%, Total time 39m 26sClaude Design, FableClaude Design, Fable: Score 8.7%, Win rate 0%, Total time 11m 50s

User-outcome comparison

Overall win rate by tool

Share of the supplied head-to-head results won by each evaluated tool.

Overall win rate100%

Task

  • Convert to Title and Content layout. Remove all projects that have X in FY25, add Halyk and Bank of South Sudan at the bottom, and fix font consistency.
Required changeSlidelyCopilotClaude DesignChatGPTClaude
Explicit tasks
Remove rows
Add rows
Change layout
Implicit tasks
Retain row markers
Follow design
Overall resultUsableUnusableUnusableUnusableUnusable
Tool output

A tool can generate a visually attractive result while changing the user’s data, deleting useful assets, ignoring the supplied PowerPoint layout, flattening an editable chart into an image, or redesigning parts of the slide that were never meant to change. The output may look good in a screenshot, yet require more work to repair than the original task would have taken manually.

At Slidely, we are building an AI agent that creates and improves complex, fully editable PowerPoint presentations. It needs to work with native layouts, placeholders, tables, charts, templates, and existing user content, not just render attractive slide images.

That requires an evaluation system designed around the way professionals actually edit presentations.

Why conventional slide evaluation falls short

Most slide evaluations focus on some combination of content quality, visual design, readability, and instruction following. These dimensions matter, but they can produce misleading results for editing tasks.

Global Performance Overview

AAcmeGlobal
Solutions

Q1 2024 Revenue by Region (USD millions)

250200150100500
210
150
120
80
North AmericaEuropeAsia PacificLatin America
USA
UK
Germany
Japan
  • Title and Content layout
  • Company logos added
  • Country flags added
  • Table converted to chart

What the PowerPoint file reveals

Global Performance Overview

Q1 2024 Revenue by Region (USD millions)

Fake layout made with
floating shapes

AAcmeGlobal
Solutions

Logos recreated with
drawing objects

Flags drawn as
shapes / emoji

250200150100500
210
150
120
80
North AmericaEuropeAsia PacificLatin America

Chart pasted as
a flat image

A slide can look correct in a screenshot but still fail as an editable, reusable presentation file.

Consider a simple request:

  • Apply the Title and Content layout.
  • Add company logos.
  • Add country flags.
  • Convert a table into a chart.

An output might appear to satisfy every instruction. But when we inspect the PowerPoint file, we may find that the layout was visually imitated using free-floating shapes, the logos were manually reconstructed using PowerPoint objects, the flags were drawn using shapes or emoji, and the chart was inserted as a flattened image.

The slide may look convincing in a screenshot. For the user, however, only the first task was completed in a usable form. The remaining elements need to be deleted and recreated. A generic visual-quality score can hide this failure.

Recent PowerPoint benchmarks have made meaningful progress beyond purely aesthetic evaluation. PPT-Eval uses task-specific rubrics, partial credit, penalties for unnecessary changes, and natural-language feedback. PPTArena focuses on in-place edits to real PowerPoint files and evaluates both visual outputs and structural differences. PresentBench shows the value of fine-grained, instance-specific evaluation criteria over broad holistic judgments.

SlidelyBench builds on this direction, but optimizes for a narrower objective: evaluating whether professional slide edits are genuinely usable.

Building SlidelyBench around real editing work

The tasks resemble the instructions people give presentation specialists and AI agents in practice:

Convert this slide to the new template.
Add the company logos and country flags.
Turn this table into a chart.
Improve the visual hierarchy without changing the content.
Remove projects marked X and add these two rows.
Fix font consistency.
Redesign this slide, but preserve all the existing assets.

These requests are often short and underspecified. There may be many good visual solutions, but there are still clear boundaries around what the tool was asked to change and what it should preserve.

The benchmark therefore does not compare outputs against one canonical screenshot. Instead, it constructs a task-specific rubric from the original slide, the instruction, and the supplied design system.

The rubric is locked before any outputs are scored. This prevents the evaluator from retrofitting criteria around whichever output looks best.

Turn every instruction into explicit and implicit tasks

Slide-editing instructions are usually short, but the requirements for a usable result are not.

A user may explicitly ask to change a layout, add an asset, or convert an object. They rarely specify every implementation detail: use the native PowerPoint layout, keep the object editable, preserve unaffected content, retain existing assets, and follow the supplied template.

A benchmark that evaluates only the literal wording of the instruction can therefore reward outputs that appear correct but are unusable.

SlidelyBench converts each request into:

  • Explicit tasks: The changes directly requested by the user.
  • Implicit tasks: The structural, preservation, and usability requirements necessary to complete those changes correctly.

Together, these form a task-specific rubric that reflects what the user actually needs, not just what the instruction literally says.

Prompt

Convert to Title and Content layout. Remove all projects that have X in FY25, add Halyk and Bank of South Sudan at the bottom, and fix font consistency.

=

Explicit tasks

  1. 1Apply the Title and Content layout
  2. 2Remove rows marked X in FY25
  3. 3Add Halyk and Bank of South Sudan
  4. 4Fix font consistency
+

Implicit tasks

  1. 1Preserve unaffected rows and slide meaning
  2. 2Retain the user’s icons and legend
  3. 3Recalculate the total after row changes
  4. 4Match the supplied template and formatting
Prompt = Explicit tasks + Implicit tasks

Score explicit tasks on usability and quality

Each explicit task is evaluated on two separate questions:

  • Was the requested result delivered in a usable form?
  • How well was it executed?

A result can look polished without being usable. For example, a table recreated with individual lines and text boxes may resemble a real table, but the user would need to delete it and rebuild it as a native PowerPoint table.

The reverse is also possible. A tool may create the correct editable object but execute the task imperfectly—for example, by omitting a requested row, using an incorrect label, or applying weak formatting.

SlidelyBench therefore assigns each explicit task:

  • Task completion: whether the requested result is fundamentally correct and usable.
  • Task quality: how accurately and professionally the completed result was executed.

A task is complete when the user can keep the delivered result and correct any remaining issues in place. It is incomplete when the result is missing, materially wrong, or must be replaced or recreated.

Each task contributes:

Task weight × task completion × task quality

This prevents visually convincing but unusable outputs from receiving credit, while still distinguishing between usable results of different quality.

Treat preservation of the user’s work as part of correctness

Completing the explicit tasks is not enough if the tool damages parts of the slide the user did not ask it to change.

SlidelyBench therefore evaluates the implicit requirements separately. It checks whether the output preserves:

  • Existing content and data.
  • Images, logos, icons, and other user assets.
  • Legends, markers, sources, and explanatory context.
  • The slide’s meaning and information structure.
  • Native PowerPoint objects and editability.
  • The supplied template and design language.

Failures are treated as unintended changes and penalized according to how much work the user must do to restore the slide:

  • Minor: a small, localized repair.
  • Medium: noticeable repair to one section.
  • Major: important content, assets, or structure must be restored.
  • Unusable: a major section, or the entire slide, effectively needs to be reconstructed.

Related changes caused by the same editing decision are grouped into a single incident rather than counted separately.

The final score is:

Final score = explicit-task score − unintended-change penalties

This ensures that a tool is rewarded not only for making the requested changes, but also for preserving everything the user still needs.

We evaluate the PowerPoint, not just the screenshot

Many properties that matter to professionals are invisible in a rendered image. Two slides can look identical while being fundamentally different PowerPoint files.

SlidelyBench checks whether:

  • The supplied slide master is actually used.
  • The correct native layout is applied.
  • The title is inside the title placeholder.
  • Source and subtitle placeholders are used when relevant.
  • Optional placeholders are allowed to remain empty.
  • Charts are native, editable PowerPoint charts.
  • Tables are native PowerPoint tables.
  • Logos and icons are real image or vector assets.
  • Theme fonts and colors are used instead of hard-coded approximations.
  • Master elements have not been recreated using slide-level shapes.
  • Objects remain sensibly grouped and editable.

We combine slide renders with PowerPoint object data rather than inferring document structure from appearance.

This reflects the product we are trying to build. Slidely works inside PowerPoint and emphasizes fully editable output, template adherence, autolayout, and spreadsheet-backed charts and tables.

Editability is not an export option added after generation. It is part of whether the task was completed correctly.

Task-specific grading creates stronger separation

A fixed visual-quality rubric asks the same broad questions for every slide. SlidelyBench changes the weighting based on the task.

For a request to add logos, flags, and a chart, the benchmark emphasizes:

  • Whether every requested asset was added.
  • Whether the assets are authentic and mapped correctly.
  • Whether the chart is native and editable.
  • Whether the source data was preserved.

For a broad redesign request, it gives more weight to information hierarchy, visual composition, readability, design-language fidelity, and improvement over the original.

For a small formatting request, it heavily penalizes unnecessary structural changes.

This makes the benchmark sensitive to the capability being tested instead of collapsing every task into a generic measure of ‘good slide design.’

In our early comparisons, this creates substantially more separation than holistic visual scoring.

Outputs that look polished but are structurally unusable or destructive fall sharply. Outputs that preserve the user’s work and make precise, PowerPoint-native edits rise.

Aligning offline evaluation with user outcomes

An offline benchmark is useful only if it predicts what users experience in the product.

Cursor describes a similar challenge in coding: an agent output can appear correct to an offline grader but feel worse to a developer using it. CursorBench combines production-derived offline tasks with controlled online evaluation to keep its definition of quality aligned with real work.

We are applying the same principle to presentations. Alongside SlidelyBench, we evaluate production signals such as:

  • Whether users accept or discard generated slides.
  • Whether they undo or regenerate an edit.
  • How much manual repair follows an agent action.
  • Whether generated objects remain in the final deck.
  • Follow-up instructions required to correct the result.
  • User preference between alternative outputs.

The strongest benchmark is not the one with the most elaborate rubric. It is the one whose rankings match the outputs professionals actually choose to keep.

What comes next

The current version of SlidelyBench focuses on [SINGLE-SLIDE / EDITING TASK DESCRIPTION]. We are expanding it toward harder workflows:

  • Compound instructions spanning several slide elements.
  • Edits across multiple slides.
  • Deck-wide consistency checks.
  • Template conversion across heterogeneous source decks.
  • Linked Excel tables and charts.
  • Multi-turn editing sessions.
  • Long-running agent workflows.
  • Consistency with organization-specific design memories.

We also plan to measure latency and compute alongside quality. The best slide agent is not simply the one that eventually produces the highest-scoring file. It must do so quickly enough to remain useful inside an interactive PowerPoint workflow.

As slide agents become more capable, visually plausible output will become easier to produce.

The harder problem is reliable editing: making exactly the requested changes, preserving everything the user still needs, and leaving behind a PowerPoint file that a professional can continue working with.

That is what SlidelyBench is designed to measure.

Get started with free credits.

Build on-brand PowerPoint decks faster than ever. Once you see the difference, you'll never go back.