We ran 18 tasks with verifiable goals through Albert and through the five most capable models available at high effort: GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, Gemini 3.1 Pro and Grok 4.6. Every task below is one Albert won on at least two of three: cheaper, faster, better outcome.
Each row shows what was asked, what Albert did, and the toughest of the five frontier models on that task: the one with the best combined price and speed among those that reached the goal. Albert was cheaper than all five on 11 of these 18 tasks and faster than all five on 18.
| Task | Albert | Toughest frontier rival | Albert's edge | Why Albert won |
|---|---|---|---|---|
Merge overlapping intervals Write a Python function that merges overlapping and touching integer intervals in O(n log n). Judged by hidden unit tests including a 200,000-interval timing check. |
GPT-5.6 Terra (medium)level 1 · score 22 $0.0063.0sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.0095.2sgoal reached |
1.31× cheaper 1.7× faster both reached the goal · beat 4 of 5 frontier models |
Albert rated the request 22/100, started GPT-5.6 Terra (medium) before the rating even came back, and returned passing code in 3.0s. The frontier models pass too, at 1.31× the price. |
Arithmetic expression evaluator Evaluate expressions with precedence, parentheses, unary minus and right-associative exponents, rejecting invalid input, without using eval. Hidden unit tests. |
Claude Sonnet 5 (medium)level 2 · score 42 $0.0415.9sgoal reached |
Claude Fable 5.1 (high)toughest of five frontier models $0.0817.2sgoal reached |
1.9× cheaper 1.09× faster both reached the goal · beat 4 of 5 frontier models |
Albert judged this harder (42/100) and escalated to Claude Sonnet 5 (medium): a bigger model only because the task earned it. It still beat the frontier models on price and speed while passing every test. |
Luhn checksum and check digit Validate card numbers with the Luhn algorithm and compute the check digit that makes a partial number valid, rejecting non-digit input. Hidden unit tests. |
GPT-5.6 Terra (medium)level 1 · score 26 $0.0098.7sgoal reached |
Claude Fable 5.1 (high)toughest of five frontier models $0.038.9sgoal reached |
3.5× cheaper 1.03× faster both reached the goal · beat 5 of 5 frontier models |
A classic that trips models on the doubling rule and spacing. Albert routed it to GPT-5.6 Terra (medium) at level 1 and passed first time, 3.5× cheaper than the toughest frontier model. |
Business days between dates Count weekdays between two ISO dates excluding a holiday list, inclusive start and exclusive end, rejecting reversed ranges. Checked against a reference implementation. |
GPT-5.6 Terra (medium)level 1 · score 15 $0.0041.7sgoal reached |
Grok 4.6 (high)toughest of five frontier models $0.0065.3sgoal reached |
1.6× cheaper 3.2× faster both reached the goal · beat 5 of 5 frontier models |
Calendar edge cases, weekends and holidays are where cheap models usually slip. Albert's worker passed every case, and the speculative start put the answer back in 1.7s. |
Topological sort with cycle reporting Order a dependency graph, include target-only nodes, and on a cycle raise an error naming the cycle. Hidden tests verify edge order and the reported cycle. |
GPT-5.6 Terra (medium)level 1 · score 35 $0.0052.4sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.0313.9sgoal reached |
6.5× cheaper 5.7× faster both reached the goal · beat 5 of 5 frontier models |
Albert scored it 35/100 and stayed at level 1: correct order, correct cycle message, 2.4s, 6.5× cheaper than the toughest frontier model. |
Roman numerals with strict validation Convert integers to Roman numerals and back, rejecting every non-canonical form such as IIII, VX or IL. Tested across all 3,999 values plus invalid inputs. |
GPT-5.6 Terra (medium)level 1 · score 34 $0.0073.9sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.0315.9sgoal reached |
3.5× cheaper 4.1× faster both reached the goal · beat 5 of 5 frontier models |
Strict validation is the trap here. Albert's level-1 worker got all 3,999 round trips and every rejection right, at a fraction of the frontier price. |
Order totals with banker's rounding Net totals per customer from a CSV, ignoring cancelled rows, rounded once to cents with round-half-to-even and no binary floating point. Exact comparison. |
GPT-5.6 Terra (medium)level 1 · score 30 $0.0082.9sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.0103.0sgoal reached |
1.15× cheaper 1.02× faster both reached the goal · beat 5 of 5 frontier models |
Banker's rounding on exact decimals is a detail most models get wrong. Albert's worker got every customer to the cent, first time, in 2.9s. |
Meeting time across daylight-saving changes Convert 09:30 New York on 8 March 2026 to London and Sydney, honouring the different daylight-saving dates. Checked against the correct times and date. |
GPT-5.6 Terra (medium)level 1 · score 18 $0.0083.2sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.028.0sgoal reached |
2.2× cheaper 2.5× faster both reached the goal · beat 5 of 5 frontier models |
A daylight-saving trap across three zones. Albert answered correctly in 3.2s for $0.008; the toughest frontier model needed 8.0s and $0.02 for the same result. |
Inclusion-exclusion count How many integers up to 1000 are divisible by 3 or 5 but not by 7, with the steps shown. Checked against the exact answer, 401. |
GPT-5.6 Terra (medium)level 1 · score 10 $0.0083.3sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.015.8sgoal reached |
1.7× cheaper 1.8× faster both reached the goal · beat 5 of 5 frontier models |
Albert rated it easy (10/100), used GPT-5.6 Terra (medium), and returned the exact answer with the working in 3.3s. |
Limerick under hard constraints Exactly five lines, every line starting with S, no line over 60 characters, must mention a cat. Checked line by line. |
GPT-5.6 Terra (medium)level 1 · score 15 $0.0084.8sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.016.7sgoal reached |
1.33× cheaper 1.41× faster both reached the goal · beat 5 of 5 frontier models |
Format constraints are checked mechanically, so there is no partial credit. Albert's worker satisfied all four rules first time; the frontier models did too at 1.33× the price. |
JSON document to a strict schema Produce a task record that validates against a strict schema with patterns, enums, array bounds and no extra keys, populated from a scenario. Validated field by field. |
GPT-5.6 Terra (medium)level 1 · score 15 $0.0071.3sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.0082.9sgoal reached |
1.13× cheaper 2.3× faster both reached the goal · beat 3 of 5 frontier models |
Albert returned a schema-perfect document in 1.3s. Speed is the story here: the toughest frontier model took 2.9s for the same output. |
Rewrite under word rules Rewrite a paragraph as exactly three sentences under 20 words each, with no word ending in -ly and every fact kept. Checked mechanically. |
GPT-5.6 Terra (medium)level 1 · score 30 $0.0094.5sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.0415.9sgoal reached |
4.3× cheaper 3.5× faster both reached the goal · beat 5 of 5 frontier models |
Every rule is checked by code, and one adverb fails the whole answer. Albert's worker met all of them first time, 4.3× cheaper and 3.5× faster than the toughest frontier model. |
Refund total from a 300-line ledger From 300 transaction lines, total the refunds for one customer excluding cancelled rows, in an exact format. Checked against computed truth. |
GPT-5.6 Terra (medium)level 1 · score 26 $0.014.4sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.068.0sgoal reached |
5.7× cheaper 1.8× faster both reached the goal · beat 5 of 5 frontier models |
Long-context arithmetic without tools. Albert's level-1 worker found all the right lines and summed them exactly; the frontier models did too at 5.7× the price. |
Current Node.js release and LTS codename Using the web, name the newest Node.js release, the newest LTS version and its codename, with a source. Checked against nodejs.org at run time. |
GPT-5.6 Terra (medium)level 1 · score 18 $0.023.6sgoal reached |
Grok 4.6 (high)toughest of five frontier models $0.0622.0sgoal reached |
2.7× cheaper 6.1× faster both reached the goal · beat 4 of 5 frontier models |
Albert recognised a request for current information and gave its worker live web search, starting the call before the routing even finished. The Claude models answered from memory and missed. |
Latest Python 3.14 patch release Using the web, give the latest 3.14 patch version and its release date. Checked against the live release data. |
GPT-5.6 Terra (medium)level 1 · score 5 $0.025.6sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.158.4sgoal reached |
7.1× cheaper 1.48× faster both reached the goal · beat 4 of 5 frontier models |
A fact newer than every model's training data. Albert searched, cited, and answered in 5.6s for $0.02; three frontier models without search got it wrong. |
Latest stable Go release Using the web, give the exact latest Go version with a source. Checked against the live release data. |
GPT-5.6 Terra (medium)level 1 · score 5 $0.023.0sgoal reached |
GPT-6 Astra (high)toughest of five frontier models $0.108.4sgoal reached |
4.8× cheaper 2.8× faster both reached the goal · beat 4 of 5 frontier models |
Same pattern: Albert turns on retrieval when the request asks for today's facts. Correct in 3.0s; the Claude models missed all three live-fact tasks. |
Bayes' theorem, worked example Explain Bayes' theorem in under 200 words with a worked medical-test example that must compute to 15.4%. Word count and result checked. |
GPT-5.6 Terra (medium)level 1 · score 8 $0.0093.5sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.027.3sgoal reached |
1.8× cheaper 2.1× faster both reached the goal · beat 4 of 5 frontier models |
Right number, right length, first time. Albert's worker delivered it in 3.5s for $0.009, 1.8× cheaper than the toughest frontier model. |
Formal German business email Translate a business email into formal German, keeping every date, figure and name, then judged for fidelity and fluency. Names, figures and register checked by code first. |
GPT-5.6 Terra (medium)level 1 · score 12 $0.0072.1sgoal reached |
Claude Opus 5 (high)toughest of five frontier models $0.0096.2sgoal reached |
1.44× cheaper 2.9× faster both reached the goal · beat 5 of 5 frontier models |
Albert scored it 12/100 and used GPT-5.6 Terra (medium): every figure and date intact, formal register, natural German, in 2.1s. |
Albert is not one model. It is a conductor that decides, per request, which model earns the job, and it does the work a single call never does.
One fast assessment rates difficulty from 1 to 100 and the category of the hardest requirement, then a deterministic table picks the cheapest model that can do the job. Harder tasks escalate; easy ones never pay frontier prices.
The likely worker begins the moment the request arrives, in parallel with the assessment. If the assessment agrees, the answer is already on its way. That is why Albert is as fast as the bare model.
If a provider goes quiet, Albert starts the next model alongside it and takes the first complete answer. Outages become seconds, not minutes.
Requests for current information get live web search automatically. Every call is priced in, including the ones Albert started and did not use.
Measured on 29 September 2026 in a single session. Coding tasks are scored by hidden unit tests, format tasks by mechanical checks, live-fact tasks against the official release data fetched at run time, and the translation by fixed checks plus an independent judge model. Each contestant had the same prompt, the same output allowance, and up to three attempts with feedback; for tasks where feedback would reveal the answer, only the first attempt counts.
Single models are priced at provider list rates with no margin. Albert's price is what a customer pays: every worker, assessment, unused speculative and hedge call, plus Albert's platform margin. Times are wall-clock from request to complete answer. Frontier models ran at high effort, their strongest standard setting.
Not shown: the 0 tasks where Albert did not win two of three against the majority of the frontier field, and the mid-tier models at medium effort, which come close to Albert on price for the easiest tasks. Albert's advantage there is speed, and staying correct when a task turns hard.