مقارنة أداء نماذج الذكاء الاصطناعي 2026

Claude Opus 5 vs Fable 5 vs GPT-5.6 Sol vs Kimi K3 vs GLM-5.2: Real Benchmark Data and Cost Per Task (2026)

Five frontier models, five leaderboards — and five different winners depending on which one you look at. As of 1 August 2026, Claude Opus 5 holds the top composite score on Artificial Analysis, GPT-5.6 Sol dominates terminal and computer-use tasks, Fable 5 ranks #1 across LMArena’s text, code, and agent arenas, and Kimi K3 leads frontend web development. Every vendor claims to be ahead. None of them are lying — they’re just measuring different things.

But here’s what the leaderboards almost never tell you: the number that should drive your decision is cost per task, not cost per token. A model that burns nearly twice as many output tokens can cost you the same total bill as a premium alternative, even when its per-token sticker price looks like a bargain. This article unpacks all of that — the published data, our own head-to-head run, and a practical guide to running your own bake-off before you commit to any provider.

Methodology

What these numbers are and where they come from

Published third-party data: Benchmark scores in the tables below are drawn from Artificial Analysis (independent model evaluation lab covering 189+ models) and from the Kylon benchmark roundup (July 2026). Pricing figures are published API list prices as of 1 August 2026. We have not modified or extrapolated any of these figures.

First-party measured run: The head-to-head section (Section 4) reports results from our own controlled test — Claude Opus 5 and Claude Fable 5 on three identical tasks, no tools, same prompt. This is a small 3-task run, not a full benchmark. We report it because it captures latency and qualitative style differences that published leaderboards miss.

GLM-5.2: Comprehensive third-party benchmarks for GLM-5.2 have not yet been publicly published as of this writing. We report only what is available and flag clearly where data is absent.

Caveat: All figures are as of 1 August 2026. Leaderboards shift weekly — a meaningful score change can appear between the time you read this and the time you deploy.

What the Published Benchmarks Say

The five models compete in genuinely different arenas. Here is what the current data shows across the key benchmarks.

Artificial Analysis Intelligence Index (Composite, 189 models)

ModelIndex ScorePercentile
Claude Opus 56098%
Claude Fable 5 (max)6098%
GPT-5.6 Sol (max)5998%
Kimi K357.197%
GLM-5.2No published score

Artificial Analysis notes that Claude Opus 5 is narrowly the most intelligent model on this composite, with comparable intelligence to Fable 5 at 26% lower cost per task. The top three are separated by only 1 index point — statistically a near-tie on a composite metric that averages dozens of tasks.

SWE-bench Verified (Software Engineering)

ModelScoreNotes
Claude Opus 596.0%Leader
Claude Fable 595.0%Close second
GLM-5.2~77.8%Provisional / self-reported
Kimi K360.4%
GPT-5.6 SolNo published scoreNot in verified row

GPQA Diamond (Graduate-Level Science)

ModelScore
GPT-5.6 Sol (max)94.1%
Kimi K393.5%
Claude Opus 5No published score in these rows

Agentic Elo (GDPval-AA v2 and AA-Briefcase)

LeaderboardLeaderScoreMargin
GDPval-AA v2 (agentic knowledge work)Claude Opus 5 (max)1861 Elo>100 pts ahead of Fable 5 and GPT-5.6 Sol
AA-BriefcaseOpus 51720 Elo+146 ahead of Fable 5

The pattern is clear: leadership is task-dependent. Opus 5 dominates agentic knowledge work and software engineering. GPT-5.6 Sol leads terminal operations, computer-use tasks (OSWorld), and graduate-level science. Fable 5 tops LMArena’s human-preference arenas. Kimi K3 leads frontend web development and BrowseComp is close (91.2 vs Sol’s 92.2). GLM-5.2 lacks comprehensive public benchmarks — its ~77.8% SWE-bench figure carries an asterisk in the source and should be treated as provisional.

Our Own Head-to-Head: Opus 5 vs Fable 5

This section reports our own first-party measured run — three identical tasks, no tools, same prompt. It is a small controlled test, not a comprehensive benchmark. We report it for the latency and qualitative data that published leaderboards do not surface.

MetricOpus 5Fable 5
Tokens used33,93233,960
Wall-clock time64.3 s77.8 s
Tools used00
Logic puzzleCorrectCorrect
Code (10 hidden tests)10 / 1010 / 10

Task 1 — Arabic copywriting (MSA): Both produced strong, idiomatic Modern Standard Arabic. The styles were genuinely different. Opus 5 was specification-led — dense technical detail: capacity, empty weight, screw cap with rubber seal, no condensation, matte finish. Fable 5 was more sensory and situational — the cold felt on the lips at the first sip, with university, beach, and intercity travel framing. Neither model reached for clichés or English loanwords. This is a stylistic difference, not a quality gap.

Task 2 — Logic puzzle: The puzzle had exactly one valid solution. Both models found it, and both identified the same unlocking deduction — the step where one constraint eliminates all alternatives. No difference in correctness.

Task 3 — Code: We executed both submissions against 10 hidden test cases they had never seen: empty input, single element, touching ranges, nested ranges, unsorted and overlapping intervals, disjoint sets, negative numbers, point ranges, a gap-of-one edge case, and duplicates. We also checked whether either model mutated the caller’s list and whether both raised ValueError on invalid input mid-list. Both scored 10/10, neither mutated the list, both raised ValueError correctly.

Honest bottom line: On these three tasks the models are evenly matched on correctness. The measurable separation was latency — Opus 5 completed in 64.3 seconds, Fable 5 in 77.8 seconds, roughly 17% faster — and writing style on the Arabic task. Neither finding generalises to a global quality claim.

Sticker Price vs Real Cost

This is the finding that almost every model comparison buries, so we’ll give it the space it deserves.

API list prices (per 1M tokens, as of 1 August 2026)

ModelInput / 1MOutput / 1MContext
GLM-5.2~$1.00~$5.001M
Kimi K3$3.00 ($0.30 cached)$15.001M
GPT-5.6 Sol$5.00$30.001.05M
Claude Fable 5$10.00$50.001M+

Looking at that table, Kimi K3 looks like a steal — roughly 30% of Fable 5’s per-token price. But here’s what the per-token rate doesn’t show you: how many tokens each model actually consumes to complete an equivalent set of tasks.

Token consumption and real output cost

When Artificial Analysis ran its Intelligence Index benchmark suite, here is what each model actually generated in output tokens — and what that costs at list price:

ModelOutput tokensOutput cost (~)Cost per Index task
Kimi K3130M~$1,950
GPT-5.6 Sol70M~$2,100
Claude Fable 587M~$4,350$2.75 (with fallback)
Claude Opus 5 (max)$2.03

The headline finding: Kimi K3’s real output cost on equivalent tasks is nearly identical to GPT-5.6 Sol’s, despite costing only half as much per output token. K3 consumes roughly 1.9x more output tokens than Sol on the same workload. The lower per-token rate is real, but it is largely offset by verbosity.

Fable 5 is the most expensive model to run in practice — not because its per-token price is highest (though it is), but because it combines high per-token cost with substantial token consumption. Opus 5 achieves a cost per task of $2.03 — versus $2.75 for Fable 5 — at matching intelligence scores. That 26% per-task cost advantage is the single most useful number in this comparison for anyone building production workloads.

Claude Opus 5 also supports cache hits at $0.50 per million tokens and offers five effort settings (low, medium, high, xhigh, max), letting you tune cost-to-quality at the task level. That kind of control matters at scale.

The Weaknesses Worth Knowing

Every model has a weakness the vendor’s marketing does not lead with. Here are the ones that matter operationally:

  • Claude Opus 5 — hallucination rate: On AA-Omniscience (factual knowledge), Opus 5 actually falls below Fable 5. More importantly: while Opus 5 gains +7 percentage points in accuracy over Opus 4.8, its hallucination rate rises +14 points. The model answers more questions — including ones where it is uncertain. For workloads where a confident wrong answer is worse than no answer, this matters. Run retrieval-augmented generation rather than relying on parametric recall.
  • Kimi K3 — SWE-bench gap: K3’s 60.4% on SWE-bench Verified is a large step below the two Claude models at 95–96%. If your primary use case is autonomous software engineering, K3 is not competitive with the Anthropic tier here.
  • GPT-5.6 Sol — SWE-bench data absent: Sol has no published score in the SWE-bench Verified leaderboard row. Draw your own conclusions about why that might be.
  • GLM-5.2 — benchmark opacity: The ~77.8% SWE-bench figure is flagged as provisional and self-reported in the original source. No comprehensive independent benchmarks have been published. That is not a disqualifier for every use case, but it means you should run your own tests before committing to production use.
  • Fable 5 — cost at scale: At $4,350 in output costs for a standard benchmark suite, Fable 5 is the most expensive model to run at scale. If you are building a high-volume product, the Opus 5 cost advantage compounds quickly.

Tutorial: Run Your Own Model Bake-Off

Published benchmarks tell you how models perform on tasks designed by benchmark authors. Your tasks are different. Here is a repeatable process for running your own head-to-head before committing to a provider. This is what we did — you can replicate it in an afternoon.

Step 1: Build a fixed task set

Pick 5–10 tasks that represent your actual workload. Not hypothetical ones — pull real examples from your production logs, customer support queue, or backlog. Aim for variety: at least one reasoning task, one code task, one open-ended generation task, and one task that touches a known model weakness (factual recall, long-context, multilingual output).

Task set template — save as tasks.json:
[
  {
    "id": "task_01",
    "type": "reasoning",
    "prompt": "A train leaves City A at 08:00 travelling at 120 km/h. A second train leaves City B at 09:30 travelling toward City A at 90 km/h. The cities are 450 km apart. At what time do the trains meet, and how far from City A?",
    "reference_answer": "12:00, 480 km from City A... [your verified solution]"
  },
  {
    "id": "task_02",
    "type": "code",
    "prompt": "Write a Python function merge_intervals(intervals) that merges overlapping intervals. Input: list of [start, end] pairs. Output: sorted merged list. Do not mutate the input.",
    "hidden_tests": ["empty_list", "single_interval", "all_overlapping", "no_overlap", "negatives"]
  }
]

Step 2: Submit identically and meter cost

Send the same prompt to each model with identical parameters — same temperature, same system prompt, same max_tokens ceiling. Log input tokens, output tokens, and wall-clock time for every call. Most API clients return this in the usage object.

Cost metering snippet (Python):
import time

def run_and_meter(client, model_id, prompt):
    start = time.time()
    response = client.messages.create(
        model=model_id,
        max_tokens=2048,
        messages=[{"role": "user", "content": prompt}]
    )
    elapsed = time.time() - start
    return {
        "content": response.content[0].text,
        "input_tokens": response.usage.input_tokens,
        "output_tokens": response.usage.output_tokens,
        "latency_s": round(elapsed, 2),
        "output_cost_usd": response.usage.output_tokens / 1_000_000 * OUTPUT_PRICE_PER_MILLION
    }

Step 3: Blind the outputs before scoring

Replace model names with codes (Model A, B, C) in your output files before any human evaluates them. Knowing which model produced which output biases scoring — even when you’re trying to be objective. This is the step most informal bake-offs skip, and it invalidates many conclusions.

Blinding script:
import json, random, string

def blind_outputs(results: dict) -> tuple[dict, dict]:
    """results = {model_name: output_text}. Returns (blinded, key)."""
    labels = random.sample(list(string.ascii_uppercase), len(results))
    blinded = {}
    key = {}
    for (model, output), label in zip(results.items(), labels):
        blinded[label] = output
        key[label] = model
    return blinded, key

Step 4: Apply a consistent scoring rubric

Write your rubric before you look at any outputs. Decide in advance what a 5/5 looks like for each task type. For reasoning: is the answer correct and is the work shown? For code: does it pass the hidden tests and avoid mutating inputs? For generation: use a 3-point scale for accuracy, tone-fit, and no hallucinations — score each dimension separately.

Rubric example for a code task (score 0–10):
- Passes all visible test cases: 3 points
- Passes all hidden test cases: 3 points
- Does not mutate input: 2 points
- Raises correct exceptions on invalid input: 1 point
- Clean, readable code (no unnecessary complexity): 1 point
TOTAL: /10

Step 5: Calculate cost per task, not total cost

Divide the total API cost for each model by the number of tasks you ran. This is the number to compare — not total tokens, not per-token price. A model that costs 20% more per token but uses 35% fewer tokens is cheaper per task. Run the same calculation for latency: average wall-clock time per task, not just a single response time measurement.

Which Model Should You Pick?

Based on the current benchmarks and cost data, here is a task-based guide. For the most up-to-date recommendations see our top AI models for 2026 comparison and our guide on how to choose the right AI model for every task.

Your primary use caseBest pickWhy
Autonomous software engineering at scaleClaude Opus 596% SWE-bench, top agentic Elo, $2.03/task
Human-preference writing and conversational productsClaude Fable 5#1 LMArena text, code, and agent arenas
Graduate-level science / research automationGPT-5.6 Sol94.1% GPQA Diamond, leads terminal and OSWorld
Frontend web development, budget-sensitive workloadsKimi K3#1 LMArena frontend, lowest sticker price — but meter actual token use
High-volume production pipelines, cost-efficiencyClaude Opus 5 or Sonnet 5$2.03/task for Opus 5; $1.53/task for Sonnet 5 (max)
Cost-sensitive experimentation / early prototypingGLM-5.2Lowest per-token price; run your own benchmarks first

Access Every Frontier Model — No International Card Needed

Claude Opus 5, GPT-5.6 Sol, Kimi K3 and more — all available through Click DZ with official licences, payment in Algerian dinar (CIB, EDAHABIA, BaridiMob), instant activation within minutes, and 24/7 local support. Rated 4.9/5 by 1,200+ customers. Save up to 60% vs official prices.

Get it on Click DZ

Frequently Asked Questions

Is Kimi K3 really cheaper than GPT-5.6 Sol in practice?

Only in some scenarios. The per-token price is lower, but Kimi K3 generates roughly 1.9x more output tokens than GPT-5.6 Sol on equivalent tasks. The two models end up at nearly identical total output cost — approximately $1,950 vs $2,100 — when running the same benchmark suite. Always measure tokens consumed, not just tokens priced.

Should the hallucination rate increase in Opus 5 be a dealbreaker?

It depends entirely on your workload. The +14 point rise reflects the model answering more questions when uncertain — which raises accuracy (+7 points) but also raises confident errors. For retrieval-augmented pipelines where the model reads provided context, this matters less. For tasks relying on parametric memory — “what is the capital of X” style recall without retrieved context — it matters a lot. Use RAG for factual workloads on any frontier model.

When will comprehensive GLM-5.2 benchmarks be available?

As of 1 August 2026, no comprehensive third-party benchmarks have been published for GLM-5.2. The available figure (~77.8% SWE-bench) is flagged as provisional in its source. Check Artificial Analysis for updates — they typically add models within weeks of their API becoming publicly available.

Conclusion

The five models covered here are all genuinely capable at the frontier. The leaderboards disagree because they measure different things, and any single ranking flattens a more complex picture. The most honest summary: Opus 5 leads on agentic software engineering and delivers the best published cost per task at $2.03. Fable 5 leads on human preference. GPT-5.6 Sol leads on science and systems tasks. Kimi K3 is the frontend specialist and lowest-sticker option, though real token consumption closes the cost gap significantly. GLM-5.2 needs independent benchmarking before you can draw firm conclusions.

Run your own bake-off on your actual tasks before committing. The tutorial above gives you everything you need to do it in a few hours. If you need to access any of these models from Algeria or North Africa without an international payment card, Click DZ offers all of them with local payment options and official licences — a practical solution for a barrier that trips up a lot of developers in the region.

اترك تعليقاً