
Five identical prompts, four current Claude and ChatGPT models, one blind reviewer. GPT-6 Astra ranked first on four of five tasks, GPT-6 Sol second, and the Claude models were correct on every number but slipped on following the brief. All 20 outputs shown unedited with credits.
Claude vs ChatGPT in September 2026: on five identical tasks, blind ranked by a reviewer who did not know which model wrote what, ChatGPT's GPT-6 Astra came first on four of five and GPT-6 Sol came second on four of five. Claude Fable 5.1 and Claude Sonnet 5 took third and fourth on every task. Every model got every number right; the ranking was decided by instruction following and fabrication. Claude Sonnet 5 invented a first-person company case study in an opinion piece, and Claude Fable 5.1 named outdated default models in a sourced-facts task. On cost the picture flips: Claude Fable 5.1 used 59 credits for the five tasks, GPT-6 Astra 38, Claude Sonnet 5 15 and GPT-6 Sol 6, so the cheapest model in this round was also the second best. Both consumer plans cost $20 a month. Every prompt, every output and every credit figure is below, unedited.

| Task | 1st | 2nd | 3rd | 4th | Reviewer's reason |
|---|---|---|---|---|---|
| Opinion piece (writing) | GPT-6 Astra | GPT-6 Sol | Claude Fable 5.1 | Claude Sonnet 5 | Astra "balances persuasion with an honest hypothetical and practical counterargument". Sonnet "invents first-person company results". |
| slugify() (coding) | GPT-6 Astra | GPT-6 Sol | Claude Fable 5.1 | Claude Sonnet 5 | Astra "preserves word boundaries, handles Unicode punctuation, and includes exactly six asserts". Fable and Sonnet "split an oversized first word despite the requirement". |
| Prices with sources (facts) | GPT-6 Astra | GPT-6 Sol | Claude Sonnet 5 | Claude Fable 5.1 | Astra "most clearly distinguishes unverified historical prices from current facts". Fable "adds speculative, potentially outdated model names without supporting sources". |
| Conversion rates (data) | GPT-6 Sol | GPT-6 Astra | Claude Sonnet 5 | Claude Fable 5.1 | Sol "correct, complete, and concise". Fable "incorrectly says the three highest rates fall in the back half: week 4 has the second-highest rate". |
| Boxes of 4 and 9 (logic) | GPT-6 Astra | GPT-6 Sol | Claude Fable 5.1 | Claude Sonnet 5 | Astra "a brief, self-contained explanation of both answers". Sonnet "confuses box counts with croissant totals in its modular proof". |
Quotes are the reviewer's notes, written against the anonymous labels and decoded afterwards. The full notes are in each task section.
| If you want | Pick | Why, from this test |
|---|---|---|
| The best answer most of the time | GPT-6 Astra | First on 4 of 5 blind rankings, no invented facts, exactly what the brief asked |
| Nearly as good for a tenth of the cost | GPT-6 Sol | Second on 4 of 5, first on the data task, 6 credits for the whole round |
| Longer, more worked-through answers | Claude Sonnet 5 | Most thorough on the logic proof and the data table, but check any "example" it gives you |
| Claude's strongest reasoning model | Claude Fable 5.1 | Correct on every number, but the most expensive model here and the one that named outdated defaults |
| To stop choosing | Krater | All four models in one chat for the same $20 a month; run the same prompt through all of them at once as we did here |
We wrote five prompts covering the things people actually ask these models to do: a long-form opinion piece, a small coding task with tests, a fact-retrieval question with sources, a data analysis with a table, and a logic puzzle. We froze the prompts before running anything and did not tune them afterwards.
All five were run on 27 September 2026 between 14:40 and 14:52 UTC in Krater's Model Arena, which sends one prompt to several models at the same time. Settings were identical for every model: manual mode, default temperature, no system prompt, no attachments, no web browsing, one turn, no follow-ups. The four models are the current Claude and ChatGPT models offered on Krater at the time of the test:
| Model | Model ID as saved | Vendor |
|---|---|---|
| Claude Fable 5.1 | anthropic/claude-fable-5.1 | Anthropic |
| Claude Sonnet 5 | anthropic/claude-sonnet-5 | Anthropic |
| GPT-6 Astra | openai/gpt-6-astra | OpenAI |
| GPT-6 Sol | openai/gpt-6-sol | OpenAI |
Two things were scored separately. Task adherence (did the output do what the prompt asked, with pass, partial or fail) was recorded by Malte from the saved outputs before the blind ranking. Quality was ranked by Vitus, who received the 20 outputs as a document with labels A to D shuffled per prompt and no model names, ranked each prompt best to worst and wrote one note per prompt. The labels were decoded only after his rankings came back. Credits are the usage cost recorded on each saved response, converted at 200 credits per dollar; they add up to the 121 credits billed to the test account for the round.
What this test is not: it is not a benchmark of image generation, voice, browsing or agents, none of which were switched on, and five prompts is a small sample. It is a fair, documented head-to-head on text tasks under identical conditions, which is what the pages currently ranking for this query do not have.
| Model | Blind rank points | Adherence | Invented facts | Wrong numbers | Credits, 5 tasks | Words, 5 tasks |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 6 (1, 1, 1, 2, 1) | 4 pass, 1 partial | 0 | 0 | 38.5 | 839 |
| GPT-6 Sol | 9 (2, 2, 2, 1, 2) | 4 pass, 1 partial | 0 | 0 | 6.1 | 727 |
| Claude Fable 5.1 | 17 (3, 3, 4, 4, 3) | 4 pass, 1 partial | 0 | 0 | 59.1 | 1,371 |
| Claude Sonnet 5 | 18 (4, 4, 3, 3, 4) | 3 pass, 2 partial | 1 | 0 | 15.3 | 1,675 |
The one partial shared by all four is the pricing task: none could name the current default chat model with confidence, and all four said so, which the prompt allowed. Sonnet's second partial is the invented case study.
| Task | Claude Fable 5.1 | Claude Sonnet 5 | GPT-6 Astra | GPT-6 Sol | Round time |
|---|---|---|---|---|---|
| Opinion piece | 12.9 | 1.5 | 17.6 | 1.1 | 41 s |
| slugify() | 20.0 | 9.1 | 7.0 | 2.4 | 39 s |
| Prices with sources | 10.6 | 1.3 | 5.5 | 0.9 | 16 s |
| Conversion rates | 11.6 | 1.2 | 5.1 | 1.1 | 14 s |
| Boxes of 4 and 9 | 4.1 | 2.3 | 3.3 | 0.6 | 10 s |
| Total | 59.1 | 15.3 | 38.5 | 6.1 |
Round time is from the prompt being sent to the slowest of the four finishing, from the saved record. We did not get a reliable per-model stopwatch on every task, so we do not publish per-model seconds; in the tasks where we watched, the GPT-6 models usually finished a few seconds ahead of the Claude models, and Astra spent the longest thinking before writing on the opinion piece.
Prompt: Write a 400 word opinion piece for a company blog arguing that four day work weeks make small software teams more productive, not less. Use one concrete example, one counterargument you address honestly, and end with a single sentence takeaway. No bullet points, no headings, no clichés like 'in today's fast paced world'.
Blind ranking: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Sonnet 5.
Reviewer's note: "C balances persuasion with an honest hypothetical and practical counterargument. B invents first-person company results and mostly argues against problems at larger teams." (C was Astra, B was Sonnet.)
Word counts: Fable 418, Sonnet 376, Astra 400, Sol 405. None used bullets, headings or the banned cliché, and all four had an example, a counterargument and a one-sentence close. The difference is what the "concrete example" was. Astra wrote "consider a hypothetical team of six developers maintaining a billing product" and built the argument on mechanisms. Sonnet wrote "when we shifted our eight-person engineering team to a four-day week last year ... over six months, we shipped the same number of features as the prior six months, and our bug count actually fell." There is no such team; the model invented a result and presented it as the author's experience. For a company blog that is a liability, and it is why Sonnet is marked partial here and last in the blind ranking.
Most arguments against a four day week for software teams rest on arithmetic: five days of work minus one day equals eighty percent of the output. That math only holds if engineers produce code at a steady rate across forty hours, and anyone who has watched a small team work knows they don't.
Consider a six-person team shipping a payments integration. In a standard week, Monday morning goes to a planning meeting, a standup, and a sync with the design lead. Wednesday afternoon is lost to a sprint retrospective that drifts into a debate about linting rules. Friday, half the team is mentally checked out by two o'clock, quietly cleaning up tickets rather than tackling the migration nobody wants to start on a Friday. Compress that week to four days and something interesting happens. The retrospective gets a hard thirty-minute cap because nobody will tolerate it eating a scarcer resource. The Wednesday debate moves to a pull request comment thread where it belongs. And the migration gets started Tuesday, because there is no longer a Friday to defer it to. The team does not lose eight hours of engineering; it loses eight hours of low-value filler that was already crowding out the real work.
Small teams benefit most because they carry the least slack. On a team of six, every meeting costs a meaningful fraction of the week's total attention, so the pressure to protect focus time is immediate and shared. Nobody needs a mandate from leadership to cancel a pointless meeting when everyone can feel its cost.
The honest counterargument is coverage. If a customer's checkout breaks on the day off, someone has to answer, and pretending otherwise is how these experiments fail. The fix is unglamorous: a rotating on-call schedule, a written escalation path, and an agreement that a genuine incident on your off day earns the time back. That adds coordination overhead, and for teams supporting critical infrastructure with thin margins, it may tip the balance the other way. This is not a plan that works for every team. It works for teams whose bottleneck is sustained, uninterrupted thinking rather than continuous availability, which describes most small product teams building software.
The measurement that matters is not hours logged but features shipped, bugs closed, and how many engineers are still on the team a year from now. On all three, a well-run four day week tends to win.
Give a small software team fewer hours and they will spend the remaining ones far more deliberately.
Every founder I know worries that cutting a day from the work week means cutting a day from output. It's an intuitive fear, and it's wrong—at least for small software teams, where the real constraint isn't hours logged but attention spent well.
When we shifted our eight-person engineering team to a four-day week last year, the immediate assumption inside the company was that our sprint velocity would drop by twenty percent. It didn't. Over six months, we shipped the same number of features as the prior six months, and our bug count actually fell. The reason wasn't magic—it was compression. Knowing Friday was off the table meant standups got sharper, meetings that used to sprawl for an hour got solved in fifteen minutes over Slack, and engineers stopped treating "I'll get to it eventually" as an acceptable posture toward code review. The scarcity of time did what a hundred productivity workshops couldn't: it forced prioritization.
The honest counterargument here is that this works because our team is small and senior. A five-person team of experienced engineers can compress communication overhead in ways a two-hundred-person org cannot; when you scale up, coordination costs multiply and a shorter week might just mean more unfinished handoffs and more context lost between people who don't sit near each other, physically or organizationally. That's a real risk, and I won't pretend otherwise. Four-day weeks are not a universal productivity hack—they're a structural change that plays to the strengths of small, autonomous teams and can genuinely hurt larger, more interdependent ones.
But for teams like ours, the four-day week isn't a perk bolted onto the existing way of working—it's a forcing function that exposes how much of the traditional five-day week was padding. Debugging a hard problem takes the same number of focused hours whether you spread them over four days or five; what changes is how much slack, meeting bloat, and half-hearted afternoon coding disappears when the deadline for "done" moves up. Rested engineers also make fewer of the small mistakes that cost hours to fix later, which is its own form of productivity nobody puts on a velocity chart.
Fewer days doesn't mean less gets built—it means less time gets wasted pretending building requires more hours than it actually does.
Small software teams should judge productivity by what they ship and sustain, not by how many weekdays they occupy. A four-day workweek, with genuinely reduced hours and unchanged pay, can make those teams more productive because it forces better choices about attention. The benefit is not that engineers suddenly type faster. It is that they spend less time switching contexts, attending marginal meetings, and repairing exhausted decisions.
Software work rarely improves in direct proportion to time spent at a desk. A clear design decision can prevent days of rework; a tired shortcut can create weeks of maintenance. For a small team, where each person carries substantial knowledge, protecting judgment matters especially. An extra day away gives people room to recover, while a shorter schedule makes unnecessary coordination harder to justify.
Consider a hypothetical team of six developers maintaining a billing product. Under its five day schedule, everyone attends daily status meetings and two weekly planning sessions. Moving to four days, the team replaces status meetings with written updates, keeps one planning session, and reserves uninterrupted mornings for development. It also limits concurrent projects to two. If it then delivers the same planned improvements with fewer defects and less rework, productivity has increased despite fewer scheduled hours. The mechanism is concrete: less coordination overhead and more focused execution, not magical motivation.
The strongest objection is coverage. Customers still encounter problems on Fridays, and a small team cannot always absorb an absent specialist. That concern is real, particularly for teams running critical services. A shorter week needs staggered days off, documented systems, and explicit support ownership. If those arrangements require chronic overtime or leave customers waiting, the policy is not working. Four days should mean fewer hours, not five days of pressure squeezed into four or unpaid work quietly pushed into the weekend.
That is why adoption should start with a carefully measured trial rather than a branding exercise. Track delivery time, escaped defects, customer response times, and actual hours worked before and during the change. Ask whether people can truly disconnect outside work. Keep commitments realistic and remove work before removing a day. Better meeting habits would help any schedule, but a shorter week gives teams a compelling reason to maintain them. Success means reliable output and healthier capacity, not merely happier survey responses.
For small software teams, a well-designed four-day week can increase productivity by protecting focus rather than attendance.
Small software teams rarely lose productivity because they lack hours. More often, they lose it because work gets interrupted, decisions wait for another meeting, and tired developers ship code that someone else must repair. That is why a well-designed four-day workweek can make a small team more productive, not less.
The key is to treat the shorter week as an operating change, not a perk layered onto the same workload. A team that closes on Fridays has a reason to make Monday planning sharper, keep meetings short, and decide what can wait. With fewer available hours, vague priorities become harder to ignore. Developers get longer stretches to build, review, and test; customers get clearer commitments instead of a crowded roadmap that slips.
Consider a hypothetical six-person team maintaining a subscription app. Its support queue regularly interrupts planned feature work, while a standing Friday meeting produces little beyond status updates. The team moves to four eight-hour days, replaces that meeting with a written update, and assigns one rotating developer to handle urgent support each day. Everyone else can protect a substantial block of focused time. The team does not magically code faster. It finishes more of what it starts because fewer people switch between tickets all afternoon, and because Thursday’s release has to be small enough to test properly.
The strongest objection is real: customers may still need help on Friday, and a deadline does not shrink because the calendar does. A four-day week can fail if managers simply compress five days of meetings and expectations into four. For teams with service obligations, coverage needs an explicit plan, whether that means staggered days off, an on-call rotation, or a narrower promise about response times. A launch may also require a temporary exception. Pretending otherwise would turn a useful practice into an avoidable source of stress.
But those constraints argue for designing the week carefully, not dismissing it. Small teams have limited capacity to begin with. An extra day of availability is worth little if it leaves everyone too scattered to solve difficult problems well. The better test is not how many hours laptops are open; it is whether the team ships dependable improvements, handles urgent issues, and can sustain that pace. Try the schedule for a quarter, measure completed work and customer response, and listen to the people doing the work.
For a small software team, four focused days can deliver more than five fragmented ones.
Prompt: Write a Python function slugify(title: str) -> str that turns a blog title into a URL slug: lowercase, ASCII letters and digits only, words joined by single hyphens, no leading or trailing hyphens, accented characters converted to their base letter (so 'Café Déjà Vu' becomes 'cafe-deja-vu'), maximum 60 characters cut at a word boundary. Include 6 test cases as asserts, including an empty string and a title made only of punctuation.
Blind ranking: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Sonnet 5.
Reviewer's note: "D preserves word boundaries, handles Unicode punctuation, and includes exactly six asserts. A merges words separated by an em dash; C and B split an oversized first word despite the requirement." (D was Astra, A was Sol, C was Fable, B was Sonnet.)
All four functions run, and all four pass the prompt's own example plus our extra checks for accents, empty input, punctuation-only input and the 60-character limit. The ranking came down to the edges. Astra's version normalises with NFKD, keeps only ASCII letters and digits, and returns an empty string rather than a truncated word if the first word alone exceeds 60 characters, which is the strict reading of "cut at a word boundary". Sol treats an em dash as nothing rather than as a separator, so "Hello—World" becomes "helloworld" in its version. Fable and Sonnet hard-cut a single word longer than 60 characters, and both wrote seven asserts where the prompt asked for six. Astra's implementation is 129 words; Sonnet's, with a docstring and comments, is 400.
import re import unicodedata MAX_SLUG_LENGTH = 60 def slugify(title: str) -> str: """Convert a blog title into a URL-safe slug. - lowercase ASCII letters and digits only - words separated by single hyphens, none leading/trailing - accented characters reduced to their base letter - at most 60 characters, truncated at a word boundary """ # 1. Decompose accented characters (é -> e + combining accent) and drop # everything that isn't ASCII. normalized = unicodedata.normalize("NFKD", title) ascii_text = normalized.encode("ascii", "ignore").decode("ascii") # 2. Lowercase, turn every run of non-alphanumerics into one hyphen, # then strip hyphens from the ends. slug = re.sub(r"[^a-z0-9]+", "-", ascii_text.lower()).strip("-") # 3. Enforce the length limit without splitting a word. if len(slug) > MAX_SLUG_LENGTH: cut = slug.rfind("-", 0, MAX_SLUG_LENGTH + 1) # If the very first word is longer than the limit, hard-cut it. slug = slug[:cut] if cut > 0 else slug[:MAX_SLUG_LENGTH] slug = slug.rstrip("-") return slug # --- tests ----------------------------------------------------------------- # accented characters map to base letters assert slugify("Café Déjà Vu") == "cafe-deja-vu" # empty input assert slugify("") == "" # punctuation only -> nothing survives assert slugify("!!! ??? ... --- ***") == "" # whitespace and punctuation collapse into single hyphens, no edge hyphens assert slugify(" Hello, World! ") == "hello-world" # mixed symbols and digits assert slugify("C++ & Python 3.11: What's New?") == "c-python-3-11-what-s-new" # long title is cut at a word boundary, never mid-word, and stays <= 60 chars long_title = "The quick brown fox jumps over the lazy dog " * 2 result = slugify(long_title) assert result == "the-quick-brown-fox-jumps-over-the-lazy-dog-the-quick-brown" assert len(result) <= 60 print("all tests passed")Notes on the implementation:
unicodedata.normalize("NFKD", ...)splits characters likeéintoeplus a combining accent; encoding to ASCII witherrors="ignore"then throws the accent away, leaving the base letter.- The single regex
[^a-z0-9]+handles spaces, punctuation, and any leftover non-ASCII in one pass, and collapsing runs into one hyphen guarantees no doubled hyphens.- Truncation searches for the last hyphen at or before position 60 (
rfindwith an exclusive end of 61), so a word ending exactly at 60 characters is kept intact. If the first word alone exceeds the limit there is no boundary to cut at, so it falls back to a hard cut rather than returning an empty slug.
import re import unicodedata def slugify(title: str) -> str: """ Convert a blog title into a URL-friendly slug. - Converts accented characters to their base ASCII letters. - Lowercases everything. - Keeps only ASCII letters/digits, joining words with single hyphens. - Strips leading/trailing hyphens. - Truncates to a maximum of 60 characters, cutting at a word boundary (never splitting a word in the middle). """ if not title: return "" # Step 1: Normalize unicode and strip accents -> base ASCII letters normalized = unicodedata.normalize('NFKD', title) ascii_str = normalized.encode('ascii', 'ignore').decode('ascii') # Step 2: Lowercase ascii_str = ascii_str.lower() # Step 3: Replace any run of non-alphanumeric characters with a single hyphen slug = re.sub(r'[^a-z0-9]+', '-', ascii_str) # Step 4: Strip leading/trailing hyphens slug = slug.strip('-') # Step 5: Truncate to max 60 chars, cutting at a word boundary if len(slug) > 60: cut = slug[:60] # If character right after the cut is not a hyphen, we're in the # middle of a word -> back off to the last hyphen inside the cut. if slug[60] != '-': last_hyphen = cut.rfind('-') if last_hyphen != -1: cut = cut[:last_hyphen] slug = cut.rstrip('-') return slug # --------------------------- # Test cases # --------------------------- # 1. Empty string assert slugify("") == "" # 2. Title made only of punctuation assert slugify("!!! *** ???") == "" # 3. Simple normal title assert slugify("Hello World") == "hello-world" # 4. Accented characters assert slugify("Café Déjà Vu") == "cafe-deja-vu" # 5. Mixed alphanumeric with punctuation assert slugify("Python 3.9 Release Notes!") == "python-3-9-release-notes" # 6. Long title requiring truncation at a word boundary (<=60 chars) long_title = "This is a really really long blog title that will definitely need truncation for sure" result = slugify(long_title) assert result == "this-is-a-really-really-long-blog-title-that-will-definitely" assert len(result) <= 60 print("All tests passed!")How it works:
- Accent removal:
unicodedata.normalize('NFKD', title)decomposes accented characters into a base letter plus a combining diacritical mark. Encoding to ASCII witherrors='ignore'then drops the diacritical marks, leaving the base letters (e.g.,é→e).- Lowercasing happens after accent stripping.
- Non-alphanumeric characters (spaces, punctuation, symbols) are collapsed into single hyphens via regex.
- Leading/trailing hyphens are stripped.
- Truncation checks the character right after the 60-char cutoff — if it's not a hyphen, it means a word was cut mid-way, so the slug backs off to the last complete word (last hyphen) within the first 60 characters.
import re import unicodedata def slugify(title: str) -> str: """Create an ASCII slug of at most 60 characters, keeping whole words. Returns "" if the first word alone exceeds 60 characters. """ normalized = unicodedata.normalize("NFKD", title) unaccented = "".join( char for char in normalized if not unicodedata.combining(char) ) words = re.findall(r"[a-z0-9]+", unaccented.lower()) result = [] length = 0 for word in words: added_length = len(word) + (1 if result else 0) if length + added_length > 60: break result.append(word) length += added_length return "-".join(result) assert slugify("") == "" assert slugify("!!! ... --- ???") == "" assert slugify("Café Déjà Vu") == "cafe-deja-vu" assert slugify(" Hello—World! Python 3 ") == "hello-world-python-3" assert slugify("a" * 56 + " boundary next") == "a" * 56 assert slugify("a" * 61) == ""
import re import unicodedata def slugify(title: str) -> str: # NFKD separates most accents from their base letters. text = unicodedata.normalize("NFKD", title.casefold()) text = text.translate(str.maketrans({ "æ": "ae", "œ": "oe", "ø": "o", "ł": "l", "đ": "d", })) text = text.encode("ascii", "ignore").decode("ascii") words = re.findall(r"[a-z0-9]+", text) slug = "" for word in words: candidate = f"{slug}-{word}" if slug else word if len(candidate) > 60: break # Stop before a word that would exceed the limit. slug = candidate return slug assert slugify("") == "" assert slugify("!!! ... ???") == "" assert slugify("Café Déjà Vu") == "cafe-deja-vu" assert slugify("Hello, WORLD! 2025") == "hello-world-2025" assert slugify("alpha " * 11) == "-".join(["alpha"] * 10) assert slugify("short " + "x" * 61) == "short"
Prompt: What do Claude Pro and ChatGPT Plus each cost per month in the United States today, and what is the name of the default model each subscription gives you in the chat app? Give the source URL for each fact. If you are not certain of something, say so rather than guessing.
Blind ranking: GPT-6 Astra, GPT-6 Sol, Claude Sonnet 5, Claude Fable 5.1.
Reviewer's note: "B most clearly distinguishes unverified historical prices from current facts. None establishes today's defaults; D adds speculative, potentially outdated model names without supporting sources." (B was Astra, D was Fable.)
We checked the facts against the vendors on the test day. claude.com/pricing shows Claude Pro at $20 a month on monthly billing, or $17 a month billed annually; anthropic.com/pricing, which two models cited, redirects there. ChatGPT Plus is $20 a month on openai.com/chatgpt/pricing. All four models gave both prices correctly with a working source URL. None of the four had web access in this test, so none could verify the current default model, and all four said so. The split is in how they said it. Astra put "last known, not verified today" in its table header and explicitly declined to name a default. Sol gave the two prices and one sentence of honest refusal in 50 words. Sonnet flagged uncertainty up front and then reasoned at length. Fable named "Claude Sonnet 4" as a "late 2025" default and GPT-5 as ChatGPT's default "in August 2025", with caveats but without sources, which is the guess the prompt asked it not to make. Neither vendor pricing page states a default model, so we score all four as partial on this task.
I can't browse the web in this conversation, so I can't verify what's true "today." What follows is based on my training data, which has a cutoff, and pricing/default models change fairly often. Please treat these as "last I knew" rather than confirmed current facts, and check the linked pages.
Claude Pro (Anthropic)
- Price: $20/month when billed monthly (about $17/month if you pay $200 annually). I'm fairly confident this was still the price as of my knowledge cutoff.
- Default model in the Claude app: As of late 2025 it was Claude Sonnet 4.5. Anthropic tends to make its newest Sonnet the default, so a newer model may have replaced it since then — I'm not certain.
- Source: https://www.anthropic.com/pricing (the plan page should list the price; the default model is usually shown inside the app's model picker rather than on the pricing page).
ChatGPT Plus (OpenAI)
- Price: $20/month. I'm fairly confident this hasn't changed.
- Default model: GPT-5 became the default for all ChatGPT users in August 2025, replacing GPT-4o. I have some recollection of a GPT-5.1 update rolling out later in 2025 as the new default, but I'm less sure of that detail.
- Source: https://openai.com/chatgpt/pricing/
Things I'm not certain about
- Whether either price has changed since my training data (both companies have adjusted plans before).
- The exact current default model name for each — this is the fact most likely to be outdated.
- Whether the URLs above have been restructured; they were the pricing pages I'm aware of, but I can't confirm they resolve today.
If you want a definitive answer, the two links above (or simply opening each app's model selector on a paid account) will settle it.
I can give you my best knowledge, but I want to flag upfront: pricing and default models for both products change fairly often, and I don't have live access to current pages, so please verify before relying on this.
What I believe is currently true
Claude Pro
- Price: $20/month (billed monthly), with a discount if billed annually
- Default model in the chat app: Claude has been shifting its "default" model naming over time (e.g., various Claude 3.x and Claude 4.x releases). I am not confident which specific model is set as default right now, since Anthropic updates this periodically and sometimes ties defaults to different tiers.
- Source: https://www.anthropic.com/pricing
ChatGPT Plus
- Price: $20/month
- Default model in the chat app: OpenAI has also changed defaults over time (GPT-4o and other variants have been the default at various points, often with an auto-switching system that picks a model based on the prompt). I am not confident which exact model is the current default.
- Source: https://openai.com/chatgpt/pricing
Why I'm hedging on the model names specifically
Both companies have released multiple model updates recently, and both apps use logic that can auto-select a model version (sometimes not even shown clearly in the UI) rather than a single static "default" you can point to with certainty. My training data has a cutoff, so a model I'd name as "current default" could easily have been superseded.
What I'd recommend
To get the accurate, current answer:
- Check https://www.anthropic.com/pricing for Claude Pro's price and check the model selector inside claude.ai (top of the chat window) for the default model shown to a new/free-tier-upgraded user.
- Check https://openai.com/chatgpt/pricing for ChatGPT Plus's price and check the model dropdown inside chatgpt.com for the default.
If you tell me today's date or paste what you see in the apps' model selectors, I can help interpret it more confidently.
I don’t have live web access in this chat, so I can’t verify today’s prices or default models. Here is what I can say without guessing:
Subscription U.S. monthly price—last known, not verified today Default model in the chat app today Official source Claude Pro $20/month on monthly billing I’m not certain. Anthropic pricing ChatGPT Plus $20/month I’m not certain. OpenAI’s “What is ChatGPT Plus?” Those URLs are official references for the prices, but I haven’t checked their current contents. I’m deliberately not naming default models: subscription model availability and the app’s default selection can change, and they aren’t necessarily the same thing.
In the United States, Claude Pro costs $20/month on monthly billing (Claude pricing), and ChatGPT Plus costs $20/month (ChatGPT pricing). Taxes may be additional.
I can’t verify the current default model name in either chat app from here, so I don’t want to guess or give you an outdated name.
Prompt: Here are weekly signups and paid conversions for a SaaS trial, weeks 1 to 8: signups 310, 295, 340, 388, 402, 376, 415, 430; paid 22, 19, 27, 33, 29, 31, 38, 36. Compute the conversion rate per week to one decimal, the overall rate, the week with the best and worst rate, and say whether the trend in conversion rate (not signups) is up, down or flat with a one sentence justification. Show a table.
Blind ranking: GPT-6 Sol, GPT-6 Astra, Claude Sonnet 5, Claude Fable 5.1.
Reviewer's note: "A is correct, complete, and concise; D is very close. B repeats the calculations unnecessarily. C incorrectly says the three highest rates fall in the back half: week 4 has the second-highest rate." (A was Sol, D was Astra, B was Sonnet, C was Fable.)
Correct answers: 7.1, 6.4, 7.9, 8.5, 7.2, 8.2, 9.2 and 8.4 percent; overall 235 of 2,956, which is 7.9 percent (7.95 unrounded); best week 7, worst week 2; trend up. All four models got every one of those right. Sol did it in 119 words with a clean table. Sonnet produced the same table and then wrote out all eight divisions a second time under a "Calculations" heading. Fable's justification for "up" included a true pooled comparison (7.6 percent for weeks 1 to 4 against 8.3 percent for weeks 5 to 8) and one false claim: the three highest rates are weeks 7, 4 and 8, so they do not all fall in the back half. Small, but the prompt asked for a one-sentence justification and the sentence contained an error.
Week Signups Paid Conversion rate 1 310 22 7.1% 2 295 19 6.4% 3 340 27 7.9% 4 388 33 8.5% 5 402 29 7.2% 6 376 31 8.2% 7 415 38 9.2% 8 430 36 8.4% Total 2,956 235 7.9% (235/2,956 = 7.95%)
- Best week: Week 7 at 9.2%
- Worst week: Week 2 at 6.4%
Trend: up. The conversion rate is noisy week to week, but the pooled rate for weeks 1–4 is 7.6% (101/1,333) versus 8.3% (134/1,623) for weeks 5–8, and the three highest rates all fall in the back half.
Weekly Conversion Rate Analysis
Week Signups Paid Conversion Rate 1 310 22 7.1% 2 295 19 6.4% 3 340 27 7.9% 4 388 33 8.5% 5 402 29 7.2% 6 376 31 8.2% 7 415 38 9.2% 8 430 36 8.4% Calculations:
- Week 1: 22/310 = 7.1%
- Week 2: 19/295 = 6.4%
- Week 3: 27/340 = 7.9%
- Week 4: 33/388 = 8.5%
- Week 5: 29/402 = 7.2%
- Week 6: 31/376 = 8.2%
- Week 7: 38/415 = 9.2%
- Week 8: 36/430 = 8.4%
Summary
Overall rate: Total paid (235) / Total signups (2,956) = 7.9%
Best week: Week 7 (9.2%)
Worst week: Week 2 (6.4%)
Trend: Up — despite a dip in Week 5, conversion rate rises from 7.1% (Week 1) to a peak of 9.2% (Week 7), and the second half of the period (Weeks 5–8, avg ~8.3%) is consistently higher than the first half (Weeks 1–4, avg ~7.5%).
Conversion rate = paid conversions ÷ signups × 100.
Week Signups Paid conversions Conversion rate 1 310 22 7.1% 2 295 19 6.4% 3 340 27 7.9% 4 388 33 8.5% 5 402 29 7.2% 6 376 31 8.2% 7 415 38 9.2% 8 430 36 8.4% Overall 2,956 235 7.9%
- Best: Week 7, at 9.2%.
- Worst: Week 2, at 6.4%.
- Trend: Up — conversion rates generally increased, rising from 7.1% in week 1 to 8.4% in week 8 despite weekly fluctuations.
Week Signups Paid conversions Conversion rate 1 310 22 7.1% 2 295 19 6.4% 3 340 27 7.9% 4 388 33 8.5% 5 402 29 7.2% 6 376 31 8.2% 7 415 38 9.2% 8 430 36 8.4% Overall 2,956 235 7.9% Best: Week 7 (9.2%). Worst: Week 2 (6.4%).
Trend: Up—despite fluctuations, the conversion rate ended higher than it started and the later weeks were generally stronger.
Prompt: A bakery sells croissants in boxes of 4 and 9. What is the largest number of croissants you cannot buy exactly using any combination of these boxes, and how many such impossible numbers are there in total? Show your reasoning briefly and give the final two numbers clearly.
Blind ranking: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Sonnet 5.
Reviewer's note: "A gives a brief, self-contained explanation of both answers. B is concise but relies on unexplained formulas; D is sound but longer. C confuses box counts with croissant totals in its modular proof." (A was Astra, B was Sol, D was Fable, C was Sonnet.)
The answers are 23 and 12 (the Frobenius number for coprime 4 and 9 is 4 x 9 - 4 - 9, and the count of non-representable positive integers is (4 - 1)(9 - 1) / 2). All four got both. Astra explained it without invoking a named theorem: the smallest buyable totals with remainders 0, 1, 2, 3 modulo 4 are 0, 9, 18, 27, so the impossible totals are the smaller numbers in each remainder class, listed out, 12 of them, the largest being 23. That is a 55-word proof a reader can check. Sol quoted the two formulas in 33 words without saying where they come from. Fable stated the formulas and then verified by listing. Sonnet wrote the longest answer, 380 words, and in its residue argument used "b = 9, 18, 27" where b was defined as the number of boxes of nine, mixing box counts with croissant totals; the final numbers were still right.
Since gcd(4, 9) = 1, the Frobenius (Chicken McNugget) theorem applies.
Largest impossible number: For coprime a and b, the largest non-representable number is ab − a − b. 4·9 − 4 − 9 = 36 − 13 = 23
Count of impossible numbers: The number of non-representable positive integers is (a−1)(b−1)/2. (4−1)(9−1)/2 = 3·8/2 = 12
Verification by listing: Numbers you can make (multiples of 4 plus 0, 9, 18, or 27): 0, 4, 8, 9, 12, 13, 16, 17, 18, 20, 21, 22, 24, 25, 26, 27, 28, …
Numbers you cannot make: 1, 2, 3, 5, 6, 7, 10, 11, 14, 15, 19, 23 — that's 12 numbers.
Once 24–27 are all achievable (four consecutive numbers), adding boxes of 4 covers everything beyond.
Final answers:
- Largest impossible number: 23
- Total impossible numbers: 12
Setting Up the Problem
We need to find non-negative integer solutions to: $$4a + 9b = n$$
Since gcd(4, 9) = 1, this system has a Frobenius number (largest impossible value) and a finite count of impossible values, given by known formulas for two coprime denominators.
Finding the Largest Impossible Number (Frobenius Number)
For two coprime integers $p$ and $q$, the largest number that cannot be represented as $pa + qb$ (with $a, b \geq 0$) is: $$g(p,q) = pq - p - q$$
Here, $p = 4$, $q = 9$: $$g(4,9) = 4 \times 9 - 4 - 9 = 36 - 13 = 23$$
Verification that 23 cannot be formed: Check $b = 0,1,2$:
- $b=0$: $4a = 23$ → not integer
- $b=1$: $4a = 14$ → not integer
- $b=2$: $4a = 5$ → not integer
No solution exists — confirming 23 is impossible.
Verification that everything larger works: For $n \geq 24$, check residues mod 4 using multiples of 9:
- $n \equiv 0 \pmod 4$: use $b=0$
- $n \equiv 1 \pmod 4$: use $b = 9$ (since $9 \equiv 1$), need $n-9 \geq 0$, true for $n\geq 9$
- $n \equiv 2 \pmod 4$: use $b = 18$ (since $18 \equiv 2$), need $n - 18 \geq 0$, true for $n \geq 18$
- $n \equiv 3 \pmod 4$: use $b = 27$ (since $27 \equiv 3$), need $n - 27 \geq 0$, true for $n \geq 27$; but check $n=24,25,26$ individually:
- $n=24$: $b=0,a=6$ ✓
- $n=25$: $b=1, 4a=16, a=4$ ✓
- $n=26$: $b=2, 4a=8, a=2$ ✓
All values $n \geq 24$ are representable, confirming 23 is the largest gap.
Counting All Impossible Numbers
For two coprime integers $p, q$, the total number of non-representable positive integers is: $$N(p,q) = \frac{(p-1)(q-1)}{2}$$
With $p=4$, $q=9$: $$N(4,9) = \frac{(4-1)(9-1)}{2} = \frac{3 \times 8}{2} = 12$$
Verification by direct listing (checking $n=1$ to $23$ for representability as $4a+9b$):
Representable: 4, 8, 9, 12, 13, 16, 17, 18, 20, 21, 22, 24, 25, 26,...
Non-representable (1 to 23): $$1, 2, 3, 5, 6, 7, 10, 11, 14, 15, 19, 23$$
Counting these: 12 values — matches the formula! ✓
Final Answer
$$\boxed{\text{Largest impossible number} = 23, \quad \text{Total impossible numbers} = 12}$$
The smallest purchasable totals with remainders (0,1,2,3) modulo (4) are (0,9,18,27), respectively. Adding boxes of 4 gives every larger total with the same remainder.
So the impossible positive totals are:
- Remainder 1: (1,5)
- Remainder 2: (2,6,10,14)
- Remainder 3: (3,7,11,15,19,23)
There are (2+4+6=12) in total.
Largest impossible number: (\boxed{23})
Total impossible numbers: (\boxed{12})
Because 4 and 9 are coprime, the largest amount that cannot be made is (4\cdot9-4-9=23). The total number of impossible positive amounts is ((4-1)(9-1)/2=12).
Largest impossible number: 23 croissants. Total impossible numbers: 12.
| Plan | Price | What you get | Source, checked 27 Sep 2026 |
|---|---|---|---|
| Claude Free | $0 | Limited daily usage of the current Claude models | claude.com/pricing |
| Claude Pro | $20 a month, $17 a month billed annually | Higher limits, Claude Code, projects, connectors | claude.com/pricing |
| Claude Max | $100 (5x Pro) or $200 (20x Pro) a month | Same features, more usage | claude.com/pricing |
| ChatGPT Free | $0 | Limited access to the current default model | openai.com/chatgpt/pricing |
| ChatGPT Plus | $20 a month | Higher limits, image generation, voice, browsing | openai.com/chatgpt/pricing |
| ChatGPT Pro | $100 (5x Plus) or $200 (20x Plus) a month | GPT-6 Pro access; the $200 tier was paused for new sign-ups on 10 Sep 2026 | openai.com/chatgpt/pricing |
| Krater Pro | $20 a month, 1,500 credits | Claude Fable 5.1, Claude Sonnet 5, GPT-6 Astra, GPT-6 Sol and 350+ other models in one chat, one credit pool | krater.ai/pricing |
At the $20 tier the two apps are the same price and the model you get is the one you pay for. The cost difference shows up when you use both, which most of the people searching this query are considering: $40 a month for two subscriptions, or one $20 Krater Pro plan with both model families in one chat, where this whole five-task, four-model round cost 121 of the 1,500 monthly credits. For deeper plan detail see our Claude pricing guide and ChatGPT Plus vs Pro.
Not in this round, on these tasks, with these models. Both GPT-6 models followed the brief more precisely, wrote less, and made no unsupported claims; the two Claude models were correct on every calculation but slipped on exactly the things a careful editor checks: one invented example, one speculative fact, one wrong justifying sentence, one extra assert. If you mostly need long, worked-through explanations, Claude Sonnet 5's outputs were the most thorough here and the second cheapest, with the caveat that its "example" in Task 1 was fiction. If you need the brief followed to the letter, GPT-6 Astra was the reviewer's first pick four times out of five, and GPT-6 Sol got within a place of it for a sixth of the cost.
The honest limit of the finding: five prompts, one run each, no browsing, no images, no voice, no agents. A different prompt set could reorder the middle two. We will rerun this when either vendor ships a new default model and update the date at the top.
/compare followed by your prompt.This round used 121 credits, about 8 percent of the 1,500 credits on the $20 Pro plan.
Not in our September 2026 blind test. On five identical text tasks, ChatGPT's GPT-6 Astra ranked first on four and GPT-6 Sol second on four; Claude Fable 5.1 and Claude Sonnet 5 ranked third and fourth on every task. All four were correct on every calculation, so the difference was instruction following and one invented example, not raw ability.
On our slugify task all four wrote working code, but GPT-6 Astra was ranked first for handling word boundaries and Unicode punctuation correctly and writing exactly the six asserts asked for. Claude Fable 5.1 and Claude Sonnet 5 both wrote seven asserts and hard-cut an over-long first word. One task is a small sample for coding; Claude Code and Codex, the vendors' agentic tools, were not part of this test.
The reviewer ranked GPT-6 Astra's opinion piece first for an honest hypothetical example and a well-handled counterargument. Claude Sonnet 5 ranked last because it invented a first-person company case study with made-up results. Claude's outputs were longer across all five tasks, which some readers prefer for explanations but which was marked as a negative here.
Both are $20 a month in the United States as of 27 September 2026. Claude Pro is $17 a month if billed annually. The heavier tiers are also matched: Claude Max and ChatGPT Pro are both $100 or $200 a month, though OpenAI paused new sign-ups to the $200 tier on 10 September 2026.
In this round GPT-6 Sol used 6.1 credits for five tasks, Claude Sonnet 5 15.3, GPT-6 Astra 38.5 and Claude Fable 5.1 59.1, measured from the usage cost on each saved response in Krater. The cheapest model was also the second best in the blind ranking.
Once in 20 outputs. Claude Sonnet 5 wrote "when we shifted our eight-person engineering team to a four-day week last year" with six months of results, in a task that asked for one concrete example. No model produced a wrong number on the data or logic tasks. Claude Fable 5.1 named outdated default models with a caveat in the facts task, which we count as a speculative claim rather than a fabrication.
The prompts were frozen before the run, all four models received identical text and settings in one Model Arena round, and the outputs were given to the reviewer with shuffled letter labels and no model names. Model identities were decoded only after all five rankings were returned. Adherence, credits and timings were recorded separately from the saved conversation records.
Yes. On Krater the Claude and GPT-6 models are in the same chat, and the /compare command sends one prompt to up to four models at once, which is how this test was run. Krater Pro is $20 a month with one credit pool, the same price as either app alone.
On five identical tasks in September 2026, blind ranked, ChatGPT's GPT-6 Astra beat both Claude models four times out of five and GPT-6 Sol did it for 6 credits total. Claude got every number right and wrote the most thorough explanations, but invented one example and guessed one fact when told not to. If you want one model, pick GPT-6 Astra; if you want the best value, GPT-6 Sol; if you want to keep checking for yourself, run the same prompt through all four in one Krater round and see who follows your brief.