NEWS
Claude Fable 5 Tops Google’s New Android Bench; Gemini Lands in Fifth
Google’s updated Android Bench has Anthropic’s Claude Fable 5 leading at 84.5% and Google’s own Gemini 3.1 Pro fifth on a new cost-aware leaderboard.
Google updated its Android Bench leaderboard this month, swapping the underlying evaluation engine to the standardized Harbor framework and adding eight new AI coding models to the ranking. The new top of the board is Anthropic’s Claude Fable 5, which scored 84.5% on the 100-task Android engineering test. Google’s own Gemini 3.1 Pro Preview sits in fifth place, and the cheapest model on the board (Deepseek V4 Flash at $1.50 per run) and the most accurate (Claude Fable 5 at $133.20) are now visibly different products on a new cost-aware leaderboard.
Android Bench is Google’s specialist leaderboard for large language models working on real Android development tasks. The benchmark’s first full re-baseline since its March launch came this month, in Google’s July update to the Android Bench methodology, and it is the first time cost, latency, and accuracy have all been re-measured on a single standardized harness.
The Methodology Reset to Harbor
Android Bench was introduced in March 2026, built on mini-swe-agent v1, a general-purpose benchmarking agent that Google had adapted to the specifics of Android development. Since then, the team has added cost and efficiency columns, evaluated open-weight models, and absorbed feedback from the Android developer community on what the ranking should actually measure. By July, Google says, the underlying model capabilities had moved far enough that the original methodology could not keep up.
The replacement is the Harbor framework. The framework’s open evaluation standard defines a common interface, containerized sandbox environments, and a way for any developer to run the benchmark against their preferred model and share results. Google re-ran every model on the board under the new framework to produce a fresh baseline, with a small shift in every score and historical numbers preserved in an online archive.
The change matters because the benchmark now runs the same way on Google’s infrastructure as on any third-party developer’s laptop, given the same Docker setup. That is the transparency Google has been promising since the March launch, and it is the part of the release that the new top-of-board numbers rest on.
By the numbers
- 25 models on the July 2026 leaderboard
- 8 new models added this month
- 100 Android tasks in the benchmark, each run 10 times
- 38,989 GitHub pull requests in the source pool
- $1.50 to $165.60 cost per full benchmark run
- 6.7 to 57.2 hours latency per full run

Claude Fable 5 Takes the Accuracy Crown
Anthropic’s Claude Fable 5 now sits at the top of the full Android Bench leaderboard with all 25 models with a score of 84.5% and a confidence interval of 79.9% to 88.8%. OpenAI’s GPT 5.5 is second at 80.2% (73.5% to 86.6%), and Anthropic’s Claude Sonnet 5 is third at 76.2% (69.0% to 82.1%). OpenAI’s GPT 5.4 holds fourth at 74.1%, and Google’s Gemini 3.1 Pro Preview rounds out the top five at 73.7%.
The eight new models added this month are Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max. Three of the five top spots go to Anthropic models, two to OpenAI, and Google’s best lands in fifth. Across the 25-model board, scores range from 84.5% at the top to 25.1% at the bottom.
Google’s blog calls Fable 5’s lead a real gap: four points above GPT 5.5. Sonnet 5’s 76.2% is in third, GPT 5.4’s 74.1% in fourth, and Gemini 3.1 Pro Preview at 73.7% in fifth. Anthropic and OpenAI hold four of the top five spots.
Claude Opus 4.8 sits in sixth at 72.4% with a confidence interval of 65.8% to 79.3%, the lowest latency in the top ten at 6.7 hours, and a $88.00 per-run cost. The leader’s score sits well above sixth place, and the rest of the board runs to 25.1% at the bottom.
The new top five, with cost and latency
| Model | Score (%) | Cost per run ($) | Latency (h) |
|---|---|---|---|
| Claude Fable 5 | 84.5 | 133.20 | 8.0 |
| GPT 5.5 | 80.2 | 138.30 | 11.4 |
| Claude Sonnet 5 | 76.2 | 99.90 | 12.3 |
| GPT 5.4 | 74.1 | 83.40 | 8.4 |
| Gemini 3.1 Pro Preview | 73.7 | 87.40 | 10.6 |
Cost figures are per full benchmark run, 100 tasks repeated 10 times. Source: Android Bench leaderboard, results as of July 8, 2026.
Gemini Lands Mid-Pack on Its Own Leaderboard
Google’s two Gemini entries on the leaderboard are mid-pack on accuracy and uneven on cost. Gemini 3.1 Pro Preview, Google’s flagship in the lineup, sits in fifth place at 73.7%, between GPT 5.4 and Claude Opus 4.8. Gemini 3.5 Flash, the smaller option, lands in eighth at 71.1%, behind GLM 5.2 from Zhipu AI and ahead of Kimi K2.7 Code.
Flash’s score is not the headline; its cost is. It is the most expensive model on the 25-model board, at $165.60 per full run, because the evaluation took 28.3 hours to complete. The 100-task, 10-run workload that most other models finished inside 12 hours stretched Flash past a full day, and the per-token charges stacked up over that runtime. By comparison, the top-scoring Claude Fable 5 finishes the same workload in 8.0 hours at $133.20.
For a benchmark designed and run by Google, the placement of Google’s own models is the awkward part. Fable 5 leads. GPT 5.5 is second. Anthropic and OpenAI take four of the top five spots, and Google’s best lands at 73.7%, well below the leader’s 79.9% lower confidence bound.
The position sits alongside a wider push inside Google, including the company’s reported effort to buy developer code for AI training, and its plan to require Gemini in software engineering interviews. A benchmark where the home team’s flagship sits in fifth is harder to spin than a generic gap.
The Cost Column Reshuffles the Field
Cost is where the new leaderboard tells a story the score column cannot. Deepseek V4 Flash is the cheapest model on the board at $1.50 per run, finishing in 8.9 hours. Its score is 54.7%, near the bottom of the accuracy range. The next cheapest, MiMo-V2.5-Pro at $9.20 per run, lands at 60.8%.
For the top of the board, cost climbs with score. Claude Fable 5 and GPT 5.5 are the most expensive models in the top five, at $133.20 and $138.30 per run. Claude Sonnet 5, third on score, costs $99.90, less than either, while Gemini 3.1 Pro Preview at $87.40 is the cheapest of the top five. Google’s flagship is competitive on price with GPT 5.4 ($83.40) but trails both on the score that decides the ranking.
The cheapest model in the top ten by cost is Kimi K2.7 Code at $48.10, ninth on accuracy at 70.4%. The mid-pack has a different optimal pick than the top of the leaderboard: smaller open-weight models look reasonable for a developer who does not need the leader’s marginal gains.
Latency reshuffles the field. Claude Opus 4.8 finishes the benchmark in 6.7 hours for $88.00, the fastest of the top ten, and Deepseek V4 Pro is third-fastest at 9.0 hours for $3.70. The slowest, Kimi K2.6, takes 57.2 hours at $49.40. The cheapest, fastest, and most accurate models are three different entries on a 25-row grid.
Score, cost, and latency for the leaders and the cheapest
| Model | Score (%) | Cost per run ($) | Latency (h) |
|---|---|---|---|
| Claude Fable 5 | 84.5 | 133.20 | 8.0 |
| GPT 5.5 | 80.2 | 138.30 | 11.4 |
| Deepseek V4 Flash | 54.7 | 1.50 | 8.9 |
| MiMo-V2.5-Pro | 60.8 | 9.20 | 13.6 |
| Qwen 3.7 Plus | 57.7 | 18.60 | 18.5 |
Source: Android Bench leaderboard, results as of July 8, 2026.
Open-Weight Models Edge Closer
The Harbor migration also brought a stronger showing from open-weight models. GLM 5.2 from Zhipu AI sits seventh overall at 72.2%, the highest-scoring open-weight model on the board. Kimi K2.7 Code from Moonshot AI follows at 70.4% in ninth. Qwen 3.7 Plus and Qwen 3.7 Max are 18th and 20th, at 57.7% and 54.2%, with Xiaomi’s MiMo-V2.5-Pro at 60.8% in 16th.
GLM 5.2 carries the second-highest latency of any model on the board at 38.9 hours per run, which lifts its cost to $117.00 despite a competitive per-token rate. Kimi K2.7 Code finishes in 31.8 hours at $48.10, and Qwen 3.7 Max in 14.2 hours at $58.30. The pattern is consistent: open-weight models at this end of the field take longer and cost less per token, which sometimes lowers the bill, sometimes raises it, depending on each model’s per-hour cost profile.
For Android developers running their own local stacks, the open-weight cluster is now scored on the same harness as the closed-source leaders, and the numbers are reproducible.
The Benchmark Opens to the Community
Along with the methodology reset, Google is letting Android developers shape the dataset itself. Through the open Android Bench project repository, anyone can submit custom Android development tasks for the team to review and, if accepted, fold into the live benchmark. Developers can also run the full evaluation against their own preferred models and publish the results.
The dataset is hosted on Harbor Hub, a public registry for Harbor-framework tasks. Google says it will review community-submitted tasks for inclusion in the live benchmark. The full dataset is 100 tasks drawn from a pool of 38,989 GitHub pull requests, with 71% written in Kotlin and 25% in Java.
Tasks in the dataset focus on Android-specific areas: Jetpack Compose for UI, Coroutines and Flows for asynchronous work, Room for persistence, Hilt for dependency injection, navigation migrations, Gradle and build configuration, and handling breaking changes across SDK updates. The task set runs from 27 lines to 435 lines of changed code per pull request, with a median of 32 lines and 46% of tasks under 27 lines.
The new contribution flow is the most operationally novel part of the release. The leaderboard was previously a one-way publication: Google ran the tests, Google published the scores. As of this release, an Android developer can write a task, run the benchmark against every model, and have the results sit on the same public page as Google’s official numbers, under the methodology Google just standardized.
What the New Columns Reveal
The Harbor reset did three things at once. It moved the methodology onto a shared, reproducible harness, re-measured every model under that harness, and added cost and latency columns that change what a developer optimizes for. The most accurate and the cheapest models are now visibly different products on the same leaderboard.
For an Android developer picking a coding model tomorrow, the practical move is to look at the score column and the cost column separately, and accept that the leaderboard is no longer answering a single question. A team with a budget can pick the accuracy leader. A solo developer running high-volume work can pick the cheapest, and still land inside the same general capability band as the mid-pack of the closed-source field.
For Google, the public scoreboard is now evidence of where its own models sit, and the team’s next move is the question the next leaderboard release will answer. The July update is the first under Harbor, and Google’s blog post invites further community contributions; the next refresh of the leaderboard will reflect what developers and the team add between now and the next re-baseline.
Frequently Asked Questions
What is Android Bench?
Android Bench is Google’s specialist leaderboard for large language models on real Android development tasks. It scores each model on 100 tasks drawn from a pool of 38,989 GitHub pull requests, runs the workload 10 times, and reports accuracy, cost, and latency. The benchmark launched in March 2026 and was rebuilt on the Harbor framework in July 2026.
What is the Harbor framework?
Harbor is a standardized evaluation framework for AI agents. It defines a common interface, containerized sandbox environments, and a way for any developer to run the same benchmark against the same model on their own infrastructure. Google adopted it for Android Bench in the July 2026 release to make the leaderboard reproducible outside its own servers.
How does Google measure cost and latency on the leaderboard?
Cost is the average per-token spend across 10 runs of the 100-task benchmark, reported in US dollars. Latency is the average wall-clock time for a full run, reported in hours. Both numbers come from the evaluation output files; Google publishes the exact figures for each model on the public leaderboard and the methodology page.
Why is Gemini 3.5 Flash the most expensive model on the board?
Flash’s per-token cost is low, but the 100-task, 10-run benchmark took it 28.3 hours, longer than any model in the top 10. The longer runtime drove the per-run bill to $165.60, more than any other model on the 25-model board. Faster models with similar per-token costs finished in 7 to 12 hours and cost less per run.
This article is based on Google’s July 2026 Android Bench release and the leaderboard data published on developer.android.com. Scores, costs, and latencies reflect the results as of July 8, 2026.
-
FINANCE2 months agoZcash Patched a Double-Spend Bug as ZEC Climbed 5%
-
ENTERTAINMENT2 months agoSteam Summer Sale 2026 Locks In June 25 to July 9 Dates
-
NEWS2 months agoMeta Adds AI Replies to Threads, But Users Can’t Block It
-
FINANCE3 weeks agoCLARITY Act Final Text Expected This Weekend as 60-Vote Hurdle Looms
-
ENTERTAINMENT2 months ago‘Widow’s Bay’ Review: Apple TV’s Sleeper Horror-Comedy Earns Its Fog
-
NEWS7 months agoFolderFresh Review: This Free Tool Automates Windows File Organizing
-
FINANCE2 weeks agoFed Minutes Cite AI Demand as Inflation Risk, Put a 2026 Hike Back on the Map
-
ENTERTAINMENT2 months agoAmazon Scraps Its Stargate Revival After a 20-Week Writers Room
