Test · September 2026

Which AI is best for presentations at a price you can actually run?

Five AI models got the same two jobs: build a profile of a well-known person from scratch, and turn a spreadsheet of monthly revenue into a six-page review. We opened every finished PowerPoint file and checked all 66 slides by hand.

The test in numbers

AI models
5
jobs each
2
presentations made
10
slides checked by hand
66
total AI cost
$3.44
total time
43 min

The short version

DeepSeek V4.1 Flash made the best slides, and cost 13 times less than the model that came second.

It scored highest on both jobs (98.5 out of 100) and used about 11 cents of AI time. Claude Sonnet 5 came second at $1.44. On the spreadsheet job four of the five models finished within a point of each other, so the real gap showed up on the open-ended job, where the model has to work out what matters on its own.

  • Paying more did not get better slides. The cheapest model won and the priciest came second.
  • Whether a deck got charts came down to the template, not the model. Changing the template took chart use from 1 deck in 6 to 5 out of 5.
  • The worst problem came from the image generator, not the writing model: it put a recognisable photo of a real person into pictures that only asked for buildings.

See all the scores →Look at the slides →

All five, scored

  1. 1DeepSeek V4.1 Flash$0.112498.5
  2. 2Claude Sonnet 5$1.437997.0
  3. 3GPT 5.6 Terra$0.562895.5
  4. 4GLM 5.2$0.536394.0
  5. 5Gemini 3.1 Pro$0.793283.0

Why these five models

This is a test of models you can afford to run on every request. If you are building a product that generates presentations for thousands of users, the model runs on every single one, so a model costing ten times more has to be ten times better before it makes sense.

Every model here sits between roughly $0.20 and $2 per million tokens of input. We deliberately left out the largest, most expensive models at the top of each vendor's range. Those would very likely score higher on the open-ended job, and they are the right choice when you are making one important deck by hand. They are the wrong choice when the bill arrives at the end of a month of production traffic.

So read the scores as “best value” rather than “best possible”. We plan to run the top-end models as a separate test later, so the two can be compared honestly: what the extra money buys, and whether it is worth it for your use case.

How the test worked

One test, set up in advance and written down before anything ran, so the results can be checked afterwards rather than taken on trust. Here is exactly what each model was given.

Job 1

A profile from scratch

Make a slides to introduce Donald Trump.

A set template, nothing else. No notes, no page count, no pictures supplied. The model has to decide what is worth saying, what to back it up with, and what to illustrate it with, which is where judgement shows up.

Job 2

A report from a spreadsheet

Create a 6 page FY26 revenue review deck for Northstar from northstar-fy26-monthly-revenue.csv in your workspace.

Nine months of made-up company revenue in a spreadsheet, an exact page count, and no template specified. Every number on every slide can be checked against the file, so download it and check ours.

Download the spreadsheetCSV · 384 bytes · 9 months · 6 columns

The exact file every model was given: monthly revenue for subscriptions, services and partners, plus customers gained and lost.

What we scored

35Followed the brief
Did it use the template it was given, hit the page count it was asked for, and actually cover the topic?
30Finished cleanly
Did it produce a working file, fix its own mistakes along the way, and avoid repeating things that were not working?
25Slide quality
How the slides look once opened in PowerPoint: readable text, nothing overlapping or cut off, sensible layout and images that fit together.
10Facts and sources
Are the claims true, and does it say where they came from? Numbers it worked out itself should be labelled as such, not passed off as reported figures.

Every point we took off has to point at a specific slide, a specific log entry or a specific line in the source file. The scores below are the average of the two jobs.

What every model got, and what we kept the same

  • Every model worked through the same presentation builder, the one inside Sharper AI. It supplies the templates, the page layouts, the chart types and the fonts, and turns whatever the model plans into a real PowerPoint file. The AI decides what goes on each slide and what it says; the builder decides how that gets drawn. So this test compares the models, not five different pieces of software.
  • All the pictures came from one image generator, OpenAI's gpt-image-2.5-flare, for every model. Nobody got a better one. That means any difference in the pictures comes from how each model described what it wanted, not from the generator itself.
  • Same settings across the board, including the same amount of room to write, so none of them were cut off early by a limit we forgot to change.
  • Three jobs ran at once, and we checked prices on the day. The costs below cover the writing model only; generating pictures is billed separately and is not included.

What these models cost

List prices on the day we ran the test, per million tokens, or very roughly 750,000 words. “Input” is what you send the model, “output” is what it writes back, and “repeat input” is the reduced rate for text the model has already seen, which most of a slide-building session is.

ModelRun throughInputper million tokensRepeat inputper million tokensOutputper million tokensCost for both decks
DeepSeek V4.1 FlashFireworks$0.22$0.022$0.66$0.1124
GLM 5.2Fireworks$1.40$0.140$4.40$0.5363
Claude Sonnet 5Amazon Bedrock$2.00$0.200$10.00$1.4379
GPT 5.6 TerraOpenAI$2.00$0.200$12.00$0.5628
Gemini 3.1 ProGoogle AI Studio$2.00$0.200$12.00$0.7932

Prices checked on 16 September 2026 and change often, so follow the links for current rates. The final column is what each model actually cost us across both jobs, which depends on how much it wrote and how much it re-read, not on list price alone.

The spread is wide: the priciest models on this list cost roughly nine times the cheapest on input and eighteen times on output. In this test the extra money bought nothing. The cheapest model scored highest overall, and the two priciest finished third and last.

All the scores

Higher scores are better. Lower cost and time are better.

ModelSlides shown belowOverall/ 100Followed the brief/ 35Finished cleanly/ 30Slide quality/ 25Facts and sources/ 10Checked in PowerPointdone / plannedAI costUS dollarsTime takensecondsActions takencountFile attemptscount
DeepSeek V4.1 Flash
fireworks2/2 checked
1198.51135.01130.01124.0119.52/211$0.1124617.1344
Claude Sonnet 5
bedrock2/2 checked
2297.01135.01130.03323.0229.02/2$1.437933554.8334
GPT 5.6 Terra
openai2/2 checked
3395.53332.51130.02223.5119.52/233$0.562811304.4334
GLM 5.2
fireworks2/2 checked
94.02234.01130.022.0338.02/222$0.5363641.1314
Gemini 3.1 Pro
google2/2 checked
83.027.52228.020.07.52/2$0.793222463.1398

All 5 models finished both jobs. Scores are the average of the two; cost and time are totals. Sorted by overall score.

Models with the same score share a place.

The slides, and why each score came out that way

Click any slide to page through the whole deck.

DeepSeek V4.1 Flash

Profile from scratch98.0

finished · saved · opened in PowerPoint · checked by hand

8 pages, more checkable facts than anyone else, and a note on every page saying where they came from. All nine pictures are objects and buildings, no people. It does credit a “Trump Presidential Library”, which does not exist.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
24 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 35/35. The fullest coverage of the five: birth date, university, taking over the family firm in 1971, the TV years, all three elections with vote counts, both impeachments, the 34-count conviction, January 6th and the 2024 shooting. None of its nine picture requests asked for a person.
  • Finished cleanly 30/30: 24 moves, three attempts at the file with two problems it fixed itself, and nothing repeated pointlessly. Cheapest of the five on this job.
  • Slide quality 24/25. We looked at all 8 finished pages: no broken words, nothing overlapping, nothing cut off. One point off because every picture is an indirect still life, so the cover tells you almost nothing about who the deck is about.
  • Facts and sources 9/10: the 2016, 2020 and 2024 vote counts all match the official results, every page names its source, and it includes unflattering facts as well as flattering ones. One point off for crediting a library that does not exist.
Report from a spreadsheet99.0

finished · saved · opened in PowerPoint · checked by hand

Exactly 6 pages, four charts with clear labels on both axes, the source file named on every page, and a closing page spelling out that October to December is missing. It got there in ten moves and one attempt: the cleanest run of the whole test.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
24 / 25
Facts and sources
10 / 10
Why this score
  • Followed the brief 35/35: the template it picked suits the numbers, it hit 6 pages exactly, every figure matches the spreadsheet, and the chart types fit what they are showing.
  • Finished cleanly 30/30: 10 moves, one attempt at the file, nothing to fix. The cleanest of all ten runs.
  • Slide quality 24/25: charts are readable with proper keys and labels. One point off because the revenue split is only ever shown as a table, never as a picture.
  • Facts and sources 10/10: we recalculated every figure and percentage against the spreadsheet and they all hold, including “net additions ranged +8 to +11 per month”. The last page states plainly what the data does not cover.

Claude Sonnet 5

Profile from scratch96.0

finished · saved · opened in PowerPoint · checked by hand

8 pages giving both terms in office equal weight, with pictures that stay generic rather than inventing anyone. On page 6 the line “FATHER OF FIVE CHILDREN.” runs over the photo frames, and no page says where anything came from.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
22 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 35/35: both terms in office covered with dates, the unusual non-consecutive precedent noted, and all eight picture requests specifically ruling out people.
  • Finished cleanly 30/30: 19 moves, two attempts at the file with one problem it fixed itself.
  • Slide quality 22/25: two points off because “FATHER OF FIVE CHILDREN.” sits on top of the photo frames near the edge of page 6; one more because page 4 carries a single line of text.
  • Facts and sources 9/10: everything it says checks out. One point off because not one of its eight pages says where any of it came from.
Report from a spreadsheet98.0

finished · saved · opened in PowerPoint · checked by hand

Exactly 6 pages, a line chart for the trend and a ring chart for the split. The quarterly customer table adds the months up correctly. It never mentions that the year is only three-quarters done.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
24 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 35/35. Six pages exactly, correct figures throughout, and the right kind of chart for each job: a line for the trend, a ring for the split. The quarterly table adds months up rather than averaging them.
  • Finished cleanly 30/30: 14 moves, two attempts at the file with one problem fixed.
  • Slide quality 24/25: clean tables and charts. One point off because three of the six pages show no data picture at all.
  • Facts and sources 9/10: every figure matches the spreadsheet. One point off for never saying that October to December is missing, which leaves “year to date” open to being read as the full year.

GPT 5.6 Terra

Profile from scratch93.0

finished · saved · opened in PowerPoint · checked by hand

6 pages and the most careful picture handling of the lot: four of its seven image requests specifically said “no people”, and nothing invented crept in. But the writing is thin: no birth date, no schooling, no election numbers.

Followed the brief
30 / 35
Finished cleanly
30 / 30
Slide quality
24 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 30/35. Five points off for how little it says: one line on page 2, four single-word cards on page 3, no birth date and no election figures. Its picture handling was the best of the five.
  • Finished cleanly 30/30: 17 moves, two attempts at the file with one problem fixed, nothing repeated pointlessly.
  • Slide quality 24/25: nothing broken anywhere and a consistent look throughout. One point off for two nearly empty pages.
  • Facts and sources 9/10: careful wording, and it says which archive things came from. One point off because it names institutions rather than pages you could actually go and read.
Report from a spreadsheet98.0

finished · saved · opened in PowerPoint · checked by hand

Exactly 6 pages, and all nine monthly totals are listed and correct. Only one chart though, and one quarter of the customer page is left blank.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
23 / 25
Facts and sources
10 / 10
Why this score
  • Followed the brief 35/35. Six pages exactly and the most complete set of numbers: all nine monthly totals listed and correct, with nothing invented.
  • Finished cleanly 30/30: 16 moves, two attempts at the file with one problem fixed.
  • Slide quality 23/25: one point off because the bottom-right quarter of page 5 is empty; one more because the monthly trend is a table rather than a chart, so the movement over time never gets shown visually.
  • Facts and sources 10/10: every figure recalculated and correct, with the period and units stated at the foot of each page.

GLM 5.2

Profile from scratch90.0

finished · saved · opened in PowerPoint · checked by hand

8 pages, accurate, and fair enough to include the 2020 defeat alongside the wins. Let down by three made-up portraits it never asked for, and a heading that breaks in the middle of a word: POLITICA / L RISE.

Followed the brief
33 / 35
Finished cleanly
30 / 30
Slide quality
20 / 25
Facts and sources
7 / 10
Why this score
  • Followed the brief 33/35: coverage is complete. Two points off because three made-up portraits went out without any note that they are not real photographs, even though the model itself only asked for buildings and atmosphere.
  • Finished cleanly 30/30: 17 moves, two attempts at the file with one problem fixed.
  • Slide quality 20/25: three points off because the page 4 heading renders as POLITICA / L RISE, split in the middle of the word, which we confirmed in the text of the finished PDF; one for thin text on that page; one because the portraits clash with the style of every other picture.
  • Facts and sources 7/10: the writing is accurate, includes the 2020 defeat and cites real sources. Three points off because the made-up portraits sit in a profile of a real person, where readers will take them for photographs.
Report from a spreadsheet98.0

finished · saved · opened in PowerPoint · checked by hand

Exactly 6 pages with two combined bar-and-line charts, and the quarterly totals add up correctly. Like Sonnet, it never says the year is incomplete.

Followed the brief
35 / 35
Finished cleanly
30 / 30
Slide quality
24 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 35/35: 6 pages exactly, with quarterly bars at 842/901/956 and the total line at 1,156/1,236/1,319, which are the correct sums.
  • Finished cleanly 30/30: 14 moves, two attempts at the file with one problem fixed.
  • Slide quality 24/25: proper keys and labels on both charts. One point off for three pages with no data picture.
  • Facts and sources 9/10: every percentage it works out is right. One point off for not saying the year is incomplete.

Gemini 3.1 Pro

Profile from scratch79.0

finished · saved · opened in PowerPoint · checked by hand

5 pages, the thinnest of the five, with no dates, schooling or numbers. The cover is a made-up photo of a real person, and the last page shows a figure with arms raised over a crowd, which reads as a tribute rather than a neutral profile.

Followed the brief
27 / 35
Finished cleanly
26 / 30
Slide quality
20 / 25
Facts and sources
6 / 10
Why this score
  • Followed the brief 27/35: six points off for saying the least of any model, with no dates, schooling or figures; two more because the cover asked for “a subtle silhouette of an American leader” and what went out is a lifelike portrait with nothing saying it is not real.
  • Finished cleanly 26/30: two points off because five of its six attempts at the file failed, the worst of the five; two more because 25 moves produced only five pages.
  • Slide quality 20/25: the layout itself is tidy. Two points off for nearly empty pages; three because the made-up cover portrait does not match anything else and the closing image reads as a tribute, which the job specifically asked us to mark against.
  • Facts and sources 6/10: what it does say is accurate. Two points off for stating flatly that the subject “permanently shifted” a party, which is an opinion presented as fact; two more for the unlabelled made-up portrait.
Report from a spreadsheet87.0

finished · saved · opened in PowerPoint · checked by hand

The chart numbers are right, but it delivered 7 pages when the job asked for exactly 6, and the extra page says only “THANK YOU”. The chart labels use the wrong units, and the nine-month total never appears anywhere.

Followed the brief
28 / 35
Finished cleanly
30 / 30
Slide quality
20 / 25
Facts and sources
9 / 10
Why this score
  • Followed the brief 28/35: five points off for handing in 7 pages when the job said exactly 6; two more because the nine-month total of $3.711M never appears.
  • Finished cleanly 30/30: 14 moves, two attempts at the file with one problem fixed.
  • Slide quality 20/25: one point off for chart labels using the wrong units; three for the empty “THANK YOU” page and for splitting one customer table over two pages under headings that do not match each other; one for having no picture beyond the two charts.
  • Facts and sources 9/10: the chart figures match the spreadsheet. One point off as the only deck in this job that never says where its numbers came from.

How the models compare, one against another

A straight AI model comparison: the same five models, ranked on the same two jobs. These are the match-ups people ask about most.

DeepSeek vs Claude

DeepSeek V4.1 Flash beat Claude Sonnet 5 on both jobs, 98.5 to 97.0, and used about 11 cents against Claude's $1.44. Claude writes slightly more polished prose and lays pages out well; DeepSeek covers more ground, names a source on every page and says plainly what its data leaves out.

Gemini vs Claude

Not close. Claude Sonnet 5 scored 97.0, Gemini 3.1 Pro 83.0, the widest gap in the test. Gemini wrote the least of any model, handed in seven pages when six were asked for, and put a made-up photo of a real person on its cover.

Claude vs GPT

Claude Sonnet 5 took 97.0 and GPT 5.6 Terra 95.5. Claude writes far more; GPT was quicker (about five minutes against nine) and handled pictures more carefully than anyone else, never inventing a person. On the spreadsheet job they finished level.

GLM vs the rest

GLM 5.2 scored 94.0 and was level with the leaders on the spreadsheet job at 98. It lost ground on the open-ended job, where the image generator added portraits it never asked for and a heading broke in the middle of a word.

What we learned

Ten presentations, 66 slides, every page looked at by hand. Four things came out of it worth knowing before you pick a model, or blame one.

  1. The cheapest model won, and it was not close on price

    DeepSeek V4.1 Flash scored 98.5 and used about 11 cents of AI time across both jobs. Claude Sonnet 5 scored 97.0 and used $1.44: 1.5 points lower for thirteen times the money. DeepSeek also named its sources on every page, included facts that do not flatter its subject, and was the only one to say which months its data left out. On this kind of work, price told us nothing useful about quality.

  2. The template decided whether you got charts, not the model

    In an earlier round with a fixed template, five of six models produced no charts at all for nine months of revenue figures. That template simply had no page layouts that could hold a chart, so any model wanting one had to abandon the house style. Letting the models pick their own template flipped it completely: all five made charts, and all five picked sensible types. If your AI slides never come back with charts, look at the template before blaming the model.

  3. The image generator added a real person nobody asked for

    GLM 5.2 asked for brick houses in Queens, the Manhattan skyline and “a rally atmosphere”. No person appears in any of its nine requests, and it specifically asked for no text in the images. The generator handed back a recognisable likeness of a living public figure on three slides. Gemini asked for “a subtle silhouette” and got a lifelike portrait. The one thing that reliably prevented this was saying “no people” outright; the models that did produced none.

  4. Simple, countable instructions still get missed

    One model handed in seven pages when the job asked for exactly six, and the extra page said nothing but “Thank you”. Another labelled its charts with the wrong units and never printed the nine-month total anywhere. These are the easiest things in the world to check and the easiest to skim past.

So which is the best AI for PowerPoint work?

Best quality for the moneyDeepSeek V4.1 FlashTop score on both jobs at a fraction of the cost, and the most careful about saying where facts came from.
When you are in a hurryGPT 5.6 TerraAbout five minutes for both decks, roughly half the field, and the most careful with images, but it writes the least.
Best writing and layoutClaude Sonnet 5The most complete, tidiest decks. Expect to pay for it, and note it never cites a source.
Turning a spreadsheet into slidesAny of the top fourOnce they could pick their own template, DeepSeek, Sonnet, GLM and Terra all landed within a point of each other.

What this test does not tell you

  • We ran each job once, not many times over. These models do not produce the same file twice, so treat this as one careful look rather than an average.
  • Two jobs, in English, with two templates. Nothing here covers editing an existing deck, other languages, or working from long documents.
  • One fixed image generator throughout, so differences in the pictures come from how each model asked, not from a better generator.
  • We make a competing tool. The scorecard, every deduction and our own mistakes running the test are published so you can check the working.
  • We only tested mid-priced models. The largest models from each vendor would likely score higher on the open-ended job, and we plan to test those separately.

Common questions

Which AI is best for making presentations?
In this test DeepSeek V4.1 Flash came out on top with 98.5 out of 100, ahead of Claude Sonnet 5 at 97.0 and GPT 5.6 Terra at 95.5. On the spreadsheet job four of the five finished within a point of each other, so the choice matters most for open-ended decks where the model has to decide what to include.
Is Claude or ChatGPT better for slides?
Claude Sonnet 5 scored slightly higher overall (97.0 against 95.5 for GPT 5.6 Terra). Claude writes considerably more and lays out pages well. GPT was about twice as fast and the most careful with images: it was the only model that specifically ruled out people in its picture requests and never produced an invented face.
Is Gemini or Claude better at presentations?
Claude, clearly. Claude Sonnet 5 scored 97.0 against 83.0 for Gemini 3.1 Pro, the biggest gap in the test. Gemini produced the thinnest decks, missed an exact page count, and put an invented photo of a real person on a cover.
Is DeepSeek any good for presentations?
It was the best of the five here, and by far the cheapest. DeepSeek V4.1 Flash won both jobs, named a source on every page, included facts that did not flatter its subject, and was the only model to state which months its data did not cover.
How much does it cost to make a presentation with AI?
Across both jobs the five models used between 11 cents and $1.44 of AI time. Ten presentations cost $3.44 in total. Generating pictures is billed separately and is not included in those figures.
Can AI build a presentation from a spreadsheet?
Yes, and this test measured exactly that. Each model got nine months of revenue in a spreadsheet and had to produce a six-page review. Four of the five got every figure right; you can download the same file and check their numbers yourself.

Related

Make one yourself

Every deck on this page came out of the presentation builder inside Sharper AI, with a different model driving each run. Give it a topic or a spreadsheet and it builds yours the same way.

Start a presentation