Benchmark · 29 September 2026 · 18 goal tasks · 5 frontier models

Albert beats the frontier models on price and speed, without giving up the goal.

We ran 18 tasks with verifiable goals through Albert and through the five most capable models available at high effort: GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, Gemini 3.1 Pro and Grok 4.6. Every task below is one Albert won on at least two of three: cheaper, faster, better outcome.

18 of 18
tasks won against the frontier field
5.2×
cheaper than GPT-6 Astra at high effort across the won tasks ($0.22 vs $1.12)
3.2×
faster than GPT-6 Astra at high effort (78s vs 248s total)
18 / 18
goals reached; the five frontier models missed 6 of 90 on the same tasks

The tasks Albert won

Each row shows what was asked, what Albert did, and the toughest of the five frontier models on that task: the one with the best combined price and speed among those that reached the goal. Albert was cheaper than all five on 11 of these 18 tasks and faster than all five on 18.

TaskAlbertToughest frontier rivalAlbert's edgeWhy Albert won
Merge overlapping intervals
Write a Python function that merges overlapping and touching integer intervals in O(n log n). Judged by hidden unit tests including a 200,000-interval timing check.
GPT-5.6 Terra (medium)level 1 · score 22
$0.0063.0sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.0095.2sgoal reached
1.31× cheaper
1.7× faster
both reached the goal · beat 4 of 5 frontier models
Albert rated the request 22/100, started GPT-5.6 Terra (medium) before the rating even came back, and returned passing code in 3.0s. The frontier models pass too, at 1.31× the price.
Arithmetic expression evaluator
Evaluate expressions with precedence, parentheses, unary minus and right-associative exponents, rejecting invalid input, without using eval. Hidden unit tests.
Claude Sonnet 5 (medium)level 2 · score 42
$0.0415.9sgoal reached
Claude Fable 5.1 (high)toughest of five frontier models
$0.0817.2sgoal reached
1.9× cheaper
1.09× faster
both reached the goal · beat 4 of 5 frontier models
Albert judged this harder (42/100) and escalated to Claude Sonnet 5 (medium): a bigger model only because the task earned it. It still beat the frontier models on price and speed while passing every test.
Luhn checksum and check digit
Validate card numbers with the Luhn algorithm and compute the check digit that makes a partial number valid, rejecting non-digit input. Hidden unit tests.
GPT-5.6 Terra (medium)level 1 · score 26
$0.0098.7sgoal reached
Claude Fable 5.1 (high)toughest of five frontier models
$0.038.9sgoal reached
3.5× cheaper
1.03× faster
both reached the goal · beat 5 of 5 frontier models
A classic that trips models on the doubling rule and spacing. Albert routed it to GPT-5.6 Terra (medium) at level 1 and passed first time, 3.5× cheaper than the toughest frontier model.
Business days between dates
Count weekdays between two ISO dates excluding a holiday list, inclusive start and exclusive end, rejecting reversed ranges. Checked against a reference implementation.
GPT-5.6 Terra (medium)level 1 · score 15
$0.0041.7sgoal reached
Grok 4.6 (high)toughest of five frontier models
$0.0065.3sgoal reached
1.6× cheaper
3.2× faster
both reached the goal · beat 5 of 5 frontier models
Calendar edge cases, weekends and holidays are where cheap models usually slip. Albert's worker passed every case, and the speculative start put the answer back in 1.7s.
Topological sort with cycle reporting
Order a dependency graph, include target-only nodes, and on a cycle raise an error naming the cycle. Hidden tests verify edge order and the reported cycle.
GPT-5.6 Terra (medium)level 1 · score 35
$0.0052.4sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.0313.9sgoal reached
6.5× cheaper
5.7× faster
both reached the goal · beat 5 of 5 frontier models
Albert scored it 35/100 and stayed at level 1: correct order, correct cycle message, 2.4s, 6.5× cheaper than the toughest frontier model.
Roman numerals with strict validation
Convert integers to Roman numerals and back, rejecting every non-canonical form such as IIII, VX or IL. Tested across all 3,999 values plus invalid inputs.
GPT-5.6 Terra (medium)level 1 · score 34
$0.0073.9sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.0315.9sgoal reached
3.5× cheaper
4.1× faster
both reached the goal · beat 5 of 5 frontier models
Strict validation is the trap here. Albert's level-1 worker got all 3,999 round trips and every rejection right, at a fraction of the frontier price.
Order totals with banker's rounding
Net totals per customer from a CSV, ignoring cancelled rows, rounded once to cents with round-half-to-even and no binary floating point. Exact comparison.
GPT-5.6 Terra (medium)level 1 · score 30
$0.0082.9sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.0103.0sgoal reached
1.15× cheaper
1.02× faster
both reached the goal · beat 5 of 5 frontier models
Banker's rounding on exact decimals is a detail most models get wrong. Albert's worker got every customer to the cent, first time, in 2.9s.
Meeting time across daylight-saving changes
Convert 09:30 New York on 8 March 2026 to London and Sydney, honouring the different daylight-saving dates. Checked against the correct times and date.
GPT-5.6 Terra (medium)level 1 · score 18
$0.0083.2sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.028.0sgoal reached
2.2× cheaper
2.5× faster
both reached the goal · beat 5 of 5 frontier models
A daylight-saving trap across three zones. Albert answered correctly in 3.2s for $0.008; the toughest frontier model needed 8.0s and $0.02 for the same result.
Inclusion-exclusion count
How many integers up to 1000 are divisible by 3 or 5 but not by 7, with the steps shown. Checked against the exact answer, 401.
GPT-5.6 Terra (medium)level 1 · score 10
$0.0083.3sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.015.8sgoal reached
1.7× cheaper
1.8× faster
both reached the goal · beat 5 of 5 frontier models
Albert rated it easy (10/100), used GPT-5.6 Terra (medium), and returned the exact answer with the working in 3.3s.
Limerick under hard constraints
Exactly five lines, every line starting with S, no line over 60 characters, must mention a cat. Checked line by line.
GPT-5.6 Terra (medium)level 1 · score 15
$0.0084.8sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.016.7sgoal reached
1.33× cheaper
1.41× faster
both reached the goal · beat 5 of 5 frontier models
Format constraints are checked mechanically, so there is no partial credit. Albert's worker satisfied all four rules first time; the frontier models did too at 1.33× the price.
JSON document to a strict schema
Produce a task record that validates against a strict schema with patterns, enums, array bounds and no extra keys, populated from a scenario. Validated field by field.
GPT-5.6 Terra (medium)level 1 · score 15
$0.0071.3sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.0082.9sgoal reached
1.13× cheaper
2.3× faster
both reached the goal · beat 3 of 5 frontier models
Albert returned a schema-perfect document in 1.3s. Speed is the story here: the toughest frontier model took 2.9s for the same output.
Rewrite under word rules
Rewrite a paragraph as exactly three sentences under 20 words each, with no word ending in -ly and every fact kept. Checked mechanically.
GPT-5.6 Terra (medium)level 1 · score 30
$0.0094.5sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.0415.9sgoal reached
4.3× cheaper
3.5× faster
both reached the goal · beat 5 of 5 frontier models
Every rule is checked by code, and one adverb fails the whole answer. Albert's worker met all of them first time, 4.3× cheaper and 3.5× faster than the toughest frontier model.
Refund total from a 300-line ledger
From 300 transaction lines, total the refunds for one customer excluding cancelled rows, in an exact format. Checked against computed truth.
GPT-5.6 Terra (medium)level 1 · score 26
$0.014.4sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.068.0sgoal reached
5.7× cheaper
1.8× faster
both reached the goal · beat 5 of 5 frontier models
Long-context arithmetic without tools. Albert's level-1 worker found all the right lines and summed them exactly; the frontier models did too at 5.7× the price.
Current Node.js release and LTS codename
Using the web, name the newest Node.js release, the newest LTS version and its codename, with a source. Checked against nodejs.org at run time.
GPT-5.6 Terra (medium)level 1 · score 18
$0.023.6sgoal reached
Grok 4.6 (high)toughest of five frontier models
$0.0622.0sgoal reached
2.7× cheaper
6.1× faster
both reached the goal · beat 4 of 5 frontier models
Albert recognised a request for current information and gave its worker live web search, starting the call before the routing even finished. The Claude models answered from memory and missed.
Latest Python 3.14 patch release
Using the web, give the latest 3.14 patch version and its release date. Checked against the live release data.
GPT-5.6 Terra (medium)level 1 · score 5
$0.025.6sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.158.4sgoal reached
7.1× cheaper
1.48× faster
both reached the goal · beat 4 of 5 frontier models
A fact newer than every model's training data. Albert searched, cited, and answered in 5.6s for $0.02; three frontier models without search got it wrong.
Latest stable Go release
Using the web, give the exact latest Go version with a source. Checked against the live release data.
GPT-5.6 Terra (medium)level 1 · score 5
$0.023.0sgoal reached
GPT-6 Astra (high)toughest of five frontier models
$0.108.4sgoal reached
4.8× cheaper
2.8× faster
both reached the goal · beat 4 of 5 frontier models
Same pattern: Albert turns on retrieval when the request asks for today's facts. Correct in 3.0s; the Claude models missed all three live-fact tasks.
Bayes' theorem, worked example
Explain Bayes' theorem in under 200 words with a worked medical-test example that must compute to 15.4%. Word count and result checked.
GPT-5.6 Terra (medium)level 1 · score 8
$0.0093.5sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.027.3sgoal reached
1.8× cheaper
2.1× faster
both reached the goal · beat 4 of 5 frontier models
Right number, right length, first time. Albert's worker delivered it in 3.5s for $0.009, 1.8× cheaper than the toughest frontier model.
Formal German business email
Translate a business email into formal German, keeping every date, figure and name, then judged for fidelity and fluency. Names, figures and register checked by code first.
GPT-5.6 Terra (medium)level 1 · score 12
$0.0072.1sgoal reached
Claude Opus 5 (high)toughest of five frontier models
$0.0096.2sgoal reached
1.44× cheaper
2.9× faster
both reached the goal · beat 5 of 5 frontier models
Albert scored it 12/100 and used GPT-5.6 Terra (medium): every figure and date intact, formal register, natural German, in 2.1s.

How Albert does it

Albert is not one model. It is a conductor that decides, per request, which model earns the job, and it does the work a single call never does.

Scores every request

One fast assessment rates difficulty from 1 to 100 and the category of the hardest requirement, then a deterministic table picks the cheapest model that can do the job. Harder tasks escalate; easy ones never pay frontier prices.

Starts before it decides

The likely worker begins the moment the request arrives, in parallel with the assessment. If the assessment agrees, the answer is already on its way. That is why Albert is as fast as the bare model.

Never waits on a stall

If a provider goes quiet, Albert starts the next model alongside it and takes the first complete answer. Outages become seconds, not minutes.

Knows when to look things up

Requests for current information get live web search automatically. Every call is priced in, including the ones Albert started and did not use.

Method

Measured on 29 September 2026 in a single session. Coding tasks are scored by hidden unit tests, format tasks by mechanical checks, live-fact tasks against the official release data fetched at run time, and the translation by fixed checks plus an independent judge model. Each contestant had the same prompt, the same output allowance, and up to three attempts with feedback; for tasks where feedback would reveal the answer, only the first attempt counts.

Single models are priced at provider list rates with no margin. Albert's price is what a customer pays: every worker, assessment, unused speculative and hedge call, plus Albert's platform margin. Times are wall-clock from request to complete answer. Frontier models ran at high effort, their strongest standard setting.

Not shown: the 0 tasks where Albert did not win two of three against the majority of the frontier field, and the mid-tier models at medium effort, which come close to Albert on price for the easiest tasks. Albert's advantage there is speed, and staying correct when a task turns hard.