Connect with us

NEWS

Flower Labs Puts Endeavor 1.0 on Customer Servers

Flower Labs launched Endeavor 1.0 a day after a £100 million UK AI contest, selling a model firms can host themselves.

Published

on

Flower Labs on 1 September 2026 launched Endeavor 1.0, a licensed generalist it says customers can run on their own servers. The London and Hamburg firm, a 2023 Cambridge spinout, is opening a production-ready preview for select organisations rather than a public download, and it is still adding compute.

The sales line is a British model that can sit next to OpenAI and Anthropic on a few public tests. The quieter move is to turn Flower’s existing federated stack, already in hospitals and banks, into a full product at the moment Whitehall started paying for home-grown AI.

Flower’s Own Scores Leave Two Tests Behind

Flower published four completed core tests against OpenAI’s GPT-5.6 Sol, Anthropic’s Claude Fable 5, Moonshot’s Kimi K3 and Nvidia’s Nemotron 3 Ultra. Endeavor records the highest HumanEval mark in that set, at 98.2, and it ties GPT-5.6 Sol and Claude Fable 5 at 99.9 on AIME 2026. It does not lead the table.

GPT-5.6 Sol is ahead on GPQA at 94.1 against Endeavor’s 92.0, a gap of 2.1 points, and ahead on IFEval at 95.9 against 94.1, a gap of 1.8. Claude Fable 5 also tops Endeavor on GPQA (92.6). Kimi K3 beats it on GPQA (93.5). Flower still says it outperforms Kimi K3 on three of the four tests and Nemotron 3 Ultra on all four. Those figures are Flower’s own launch numbers. Nobody else has rerun them in public.

LAUNCH SCORES ON FOUR PUBLIC TESTS

Test Endeavor 1.0 GPT-5.6 Sol Claude Fable 5 Kimi K3 Nemotron 3 Ultra
GPQA 92.0 94.1 92.6 93.5 86.7
HumanEval 98.2 95.1 97.0 96.3 96.3
IFEval 94.1 95.9 91.7 92.8 91.9
AIME 2026 99.9 99.9 99.9 96.7 94.2

A 98.2 HumanEval print, 1.2 points above Claude and 3.1 above GPT-5.6 Sol, is a strong coding headline and a thin basis for a “frontier-class” claim. The company says no small set of tests can describe real use, then leads with those four anyway because that is the language a buyer’s board already knows. A high coding mark also says little about whether a multi-step agent job will hold up inside a ward or a dealing room, which is the work Flower says it built the model for.

Flower’s own account, posted as the model went live, put the scores next to the hosting pitch.

https://x.com/flwrlabs/status/2094692986919817250

The datasheet linked from the model page still does not state a parameter count. Endeavor draws on “mature, widely available capabilities” from leading open-weight models, then adds UK-specialist behaviour from Flower’s earlier Lizzy model and new in-house training. That is a stacked system, not a lab that pretends it trained a rival to GPT from a blank checkpoint.

The NHS Already Trains Without Moving Files

Flower’s older product is the reason Endeavor has somewhere to land. The company grew out of an open-source framework for distributed training that Daniel J. Beutel, Taner Topal and Nicholas Lane built from Cambridge research, then took through Y Combinator’s Winter 2023 batch. Federated learning leaves the files where they sit and moves the model instead. A hospital can train on patient records that never leave the trust.

That is not a slide. In March 2025 Flower and the BloodCounts! consortium published results from a federated full blood count analysis across hospitals in the UK, the Netherlands and the Gambia. The study used Flower nodes inside each site to hunt early haematological disease, including leukaemia, without pooling raw records. Newly opened centres and sites with a high share of ethnic minority patients saw the largest lift, of over 9% in balanced accuracy on iron deficiency, a condition the consortium ties to anaemia that affects over 1.9 billion people.

THE BLOODCOUNTS HOSPITAL NETWORK

  • First sites: Hospitals in the UK, the Netherlands and the Gambia ran the first Flower-powered study.
  • NHS names: Cambridge University Hospitals NHS Trust, NHS Blood and Transplant, University College London Hospitals and Barts Health sit on the consortium list.
  • Scale plan: Twenty hospitals were due to join by the end of 2025, with the rest of the group following in 2026 toward 108 sites.
  • Other Flower users: The firm names the NHS and JP Morgan as clients, and in 2024 listed Samsung, Nokia Bell Labs, Brave and Banking Circle among early framework adopters.

Lane is co-founder and chief scientist and a professor of machine learning at Cambridge. Beutel is chief executive; Topal is chief operating officer. Felicis led a $20 million Series A on 15 February 2024, nine months after a $3.6 million pre-seed, at a $100 million valuation. Mozilla Ventures and Hugging Face chief executive Clem Delangue were among the backers. In 2026 the company says it has raised over £23 million. The federated layer is what those buyers already paid for. Endeavor is the generalist they can now hang on it.

A £100 Million Window Opened a Day Early

On 31 August 2026, one day before Endeavor shipped, the government opened the first competitions under a £100 million scheme to buy British AI for public services. Chancellor John Healey announced the Sovereign AI R&D Procurement Scheme at a G20 meeting and said Britain should back firms to start, scale and succeed in the UK. AI minister Kanishka Narayan said the unit would “put the heft of a nation behind Britain’s AI founders.”

The first four contests map onto Flower’s pitch with uncomfortable neatness. One is an NHS productivity challenge with the Department of Health and Social Care, asking firms to automate workflows, coordinate care and support decisions. Another, with the Ministry of Defence, asks for ways to connect data and frontier models across defence systems. A third, with the National Cyber Security Centre, covers agent security. A fourth targets cheaper public compute. Winners keep the intellectual property. Smaller firms can get paid up front so they are not locked out by turnover tests.

FROM FEDERATED CODE TO ENDEAVOR

  1. 8 March 2023: Flower Labs launches in public with Y Combinator after two years as an academic project.
  2. 15 February 2024: Felicis leads a $20 million Series A at a $100 million valuation.
  3. 20 March 2025: BloodCounts! and Flower publish the first federated blood-count results with NHS trusts in the mix.
  4. 15 April 2026: Flower releases Lizzy 7B, an open-weight UK-specialist model.
  5. 31 August 2026: The UK opens the £100 million Sovereign AI procurement contests, led by an NHS productivity brief.
  6. 1 September 2026: Flower launches Endeavor 1.0 as a managed or private preview.

The Sovereign AI Unit behind that scheme is backed by up to £500 million. A Turing Institute security brief this summer treated model possession and operational control as the hard end of sovereignty, the point where a contractual API is no longer enough for high-stakes work. Flower is selling exactly that control, on the first morning a British buyer had a new pot of money to spend on it. There is no public sign Flower has entered the NHS contest. The timing still does the commercial work.

THE FOUR CONTESTS OPENED ON 31 AUGUST

  • NHS productivity: Automate workflows and support decisions across health services, with the Department of Health and Social Care.
  • Compute efficiency: Cut the cost of public AI compute, with ARIA’s Scaling Inference Lab.
  • Defence integration: Connect data and frontier models across Ministry of Defence systems.
  • Agent security: Test and contain risks from capable agents, with the National Cyber Security Centre.

Healey’s line was that people in every postcode should feel better public services from British AI. Flower already has NHS trusts on its federated network. Endeavor gives those trusts a generalist they can keep inside the same walls.

Lizzy Came With Open Weights

Endeavor is the second model in Flower’s programme, four months after Lizzy 7B, which the company released on 15 April 2026 as a UK-built open-weight model. Lizzy was small enough to run locally and was posted for download, with GGUF builds for machines that will never see a datacentre. Flower said it was trained with UK institutions, language and use cases in mind, and it posted specialist “Britishness” tests beside ordinary reasoning scores.

On those tables Lizzy 7B scored 77.9 on MATH against 31.3 for EuroLLM 9B and 22.4 for Switzerland’s Apertus 8B, 89.9 on Britishness Domains against 69.0 and 32.6, and 34.6 on GPQA against 26.8 and 28.1. That is a compact specialist beating larger European sovereign peers on the tests Flower chose. It is also a different product from Endeavor. Lizzy was a 7 billion parameter assistant you could pull from Hugging Face. Endeavor is a licensed generalist you request, then run through Flower or inside your own environment.

The company is explicit about the jump. Lizzy showed that a model can carry local knowledge and still meet a deployment constraint. Endeavor, it says, is a broad generalist meant to compete with leading frontier systems on reasoning, software and long-horizon agent work. The weights are no longer the giveaway. The giveaway is a path to host the thing yourself after you have built agents and evals around it. Open-weight Lizzy was a flag. Licensed Endeavor is a contract.

Who Can Run Endeavor 1.0 Today?

Access is by request. Flower is onboarding a limited set of organisations and partners for managed and private installs, and it will widen that list as it expands compute. There is no self-service endpoint and no public weight dump. The company will operate Endeavor as a production service for customers who want it to handle scaling, and it will support private installs for teams that need the model inside their own environment, including a split in which most work stays on Flower’s service and sensitive jobs run locally.

WHAT WE KNOW

  • Access path: Preview by request, managed service or private install, with a later wider rollout tied to more compute.
  • Licence: Endeavor is sold under licence; clients can train on data that does not move to a central server.
  • Harness: Flower says it tuned how the model spends reasoning effort, keeps context, calls tools and recovers when a step fails.
  • FlowerBench: Extra signals come from opted-in enterprise tasks that run inside the customer’s own environment so proprietary files stay put.

WHAT IS UNCONFIRMED

  • Size: No parameter count, training-token figure or hardware recipe has been published for Endeavor 1.0.
  • Independent tests: The four-way table has not been replicated by a third party.
  • Named buyers: Flower has not identified the first Endeavor preview customers.
  • Public money: There is no public filing that Flower has bid into the new NHS or defence contests.

Familiar APIs and response formats are the on-ramp, so existing apps and agent frameworks can call it. Private deployment is the clause that matters for a hospital or a bank. Once agents, evals, tool wiring and data pipes grow around a model, a closed API that later changes price, policy or availability becomes a product risk. Flower is selling the option to pick up the same model and walk it inside the firewall.

The Product Is the Path Off a Closed API

Lane put the political wrapping on that clause when he said Europe should stop hiring its brains by the token.

Europe should not have to rent its intelligence indefinitely from a handful of US companies.

Professor Nicholas Lane, co-founder and chief scientist, Flower Labs

The launch post makes the same point in product language. Flower says it wants organisations to start with frontier capability and then extend the model with their own agents, evaluations, data pipelines and improvement loops, so the intelligence becomes “increasingly your own.” Managed use is the fast start. Private install is the exit. The company writes that choosing an endpoint should not have to become a permanent infrastructure decision.

That is a sharper offer than another chatbot score. It is also an unfinished one. Endeavor still needs more compute before Flower will open the door. The scores that sell the door are Flower’s own, and they already show GPT-5.6 Sol ahead on two of the four tests the company chose to publish. The NHS productivity contest is open now, and Flower already knows how to train on blood counts without moving a patient file. The preview list is where those two facts meet.

As the founder of Thunder Tiger Europe Media, Dr. Elias Thornwood brings over 25 years of experience in international journalism, having reported from conflict zones in the Middle East, Asia, and Africa for outlets like BBC World and Reuters. With a PhD in International Relations from Oxford University, his expertise lies in geopolitical analysis and global diplomacy. Elias has authored two bestselling books on European foreign policy and received the Pulitzer Prize for International Reporting in 2015, establishing his authoritativeness in the field. Committed to trustworthiness, he enforces rigorous fact-checking protocols at Thunder Tiger, ensuring unbiased, evidence-based coverage of worldwide news to empower informed global audiences.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending