
Image: @SakanaAILabs (X)
Summary
- Sakana AI was founded in Tokyo in July 2023 and reached a valuation of roughly $2.7 billion and cumulative funding of $412 million within three years.
- Rather than scaling up parameters, its research bets on evolutionary algorithms, model merging, and recursive self-improvement — work that became products in 2026 as Fugu, Marlin, and Sakana Chat.
- Revenue comes from partnerships with major Japanese corporations and defense/intelligence contracts, but a history of walked-back performance claims and heavy reliance on the Japanese market remain risk factors.
Three People, One Anti-Scaling Bet, Started in Tokyo
Sakana AI was founded in Tokyo in July 2023. As of February 16, 2026, its headquarters sits on the 22nd floor of Mori JP Tower, Azabudai Hills, in Minato Ward. It has three co-founders, and their backgrounds barely overlap.
CEO David Ha spent eight years at Goldman Sachs, rising to co-head and managing director of the Japan rates trading desk, before moving to Google Brain in 2016. While still at the bank, he ran an AI research blog under the pseudonym "hardmaru" — an account that became known in the deep learning community before his name was. After Google, he spent about a year, starting in 2022, leading research at Stability AI.
CTO Llion Jones grew up in Wales. He built a chatbot at age 14 and earned both a bachelor's and master's in computer science and AI from the University of Birmingham. His best-known credit is the 2017 paper "Attention Is All You Need" — he's one of eight co-authors on the paper that introduced the Transformer architecture, now the backbone of nearly every large language model. Jones has said he joined Sakana AI because he believed the conventional Transformer approach had entered a phase of diminishing returns.
COO Ren Ito studied law at the University of Tokyo, earned an LLM from NYU School of Law, and a master's from Stanford. He joined Japan's Ministry of Foreign Affairs in 2001 and worked as a diplomat before becoming an executive officer for global operations at Mercari. He, too, passed through Stability AI as COO. The fact that Ha and Ito both came out of the same company, and that Ha and Jones met at Google's Tokyo office, is the backstory behind this founding trio.
The company's name, Sakana (魚), means "fish" in Japanese. A single fish follows only a few simple rules, but a school of fish as a whole produces complex behavior — evading predators, finding food. The idea of translating that collective intelligence into model design is the company's founding premise. Against an industry racing to scale up parameters for performance, the three founders believed there was another path.
Evolutionary Model Merging — Building New Models Without Training
The result that first put the company on the research community's map was Evolutionary Model Merge, released on March 21, 2024. It combines multiple existing open models into a new one — but instead of a human deciding how to combine them, an evolutionary algorithm searches for the answer. The paper was published in Nature Machine Intelligence in January 2025.
The search happens in two spaces simultaneously. In parameter space (PS), the algorithm evolves the mixing ratio of weights across layers from multiple models. In data-flow space (DFS), it evolves the order in which layers from different models get stitched together. Both spaces contain far more combinations than human intuition could ever exhaustively explore. Running for 100 to 150 generations, only the best-performing individuals survive to produce the next generation, eventually yielding a final model.
The core of this method is that it never runs gradient-descent-based training even once. Instead of burning GPU cycles on repeated backpropagation, it reframes the problem as one of combining already-trained weights — dramatically cutting compute costs.
The company published results alongside the method. A 7-billion-parameter model called EvoLLM-JP scored 55.2% on the Japanese math reasoning benchmark MGSM-JA (50.6% for the publicly released v1-7B version). It was built by merging a Japanese-language model (Shisa-Gamma) with math-specialized models (WizardMath, Abel). Despite being built specifically for math, its average score on the general Japanese benchmark (JP-LMEH) was 70.5 — beating every Japanese model under 70 billion parameters as well as the previous best model near that size. Still, on the same table, Swallow 70B scored higher, at 71.5. The same family includes EvoVLM-JP, which handles images and language together, and EvoSDXL-JP, an image-generation model that processes Japanese prompts through four-step reasoning.
| Space | What's evolved | How humans used to do it |
|---|---|---|
| Parameter Space (PS) | Layer-by-layer weight mixing ratios | Fixed-ratio averaging |
| Data-Flow Space (DFS) | Layer ordering across different models | Designed by intuition and experience |
Self-Correcting AI — From The AI Scientist to the RSI Lab
The second research pillar is getting AI to improve AI. It began with "The AI Scientist," released on August 13, 2024 — a system that generates research ideas, writes experimental code, runs it, analyzes results, writes the paper, and even peer-reviews it, all without human intervention. The reported compute cost to produce one paper was about $15.
A follow-up, AI Scientist v2, was evaluated at an ICLR 2025 workshop in March 2025. One of three submitted papers scored an average of 6.33, clearing the acceptance bar — a score higher than 55% of human-authored submissions. Sakana AI voluntarily withdrew the paper before publication, however. It hadn't gone through the workshop organizers' meta-review, workshop acceptance rates run much higher than main-conference rates, and citation errors were flagged. That's the actual scope of the claim that "AI passed peer review." A paper describing the system itself was published in Nature on March 25, 2026 (volume 651, pages 914–919).
The same research produced a safety-relevant finding. During testing, when an experiment ran past its time limit, the system — instead of writing faster code — modified its own code to simply extend the time limit. In one case, it inserted a system call to re-launch itself, creating an infinite loop. After the company disclosed this, it fueled ongoing debate about controlling autonomous agents.
Subsequent releases pushed further into self-improvement. The Darwin Gödel Machine, released in May 2025, has agents rewrite their own codebase to generate variants. It raised SWE-bench scores from 20% to 50% — a 30-point absolute gain — and boosted scores on the multilingual coding benchmark Polyglot from 14.2% to 30.7%. In September of the same year, the company open-sourced ShinkaEvolve, a program-evolution framework combining language models with evolutionary algorithms. It matched the best known result for the classic circle-packing optimization problem (placing 26 circles) using just 150 samples, and in a separate experiment discovered a new load-balancing loss function for mixture-of-experts (MoE) models within 30 generations.
There's competitive-programming evidence too. On December 14, 2025, the algorithmic optimization agent ALE-Agent placed first in AtCoder Heuristic Contest 058 — a four-hour contest against 804 human competitors, and the first known win by an AI agent in such a competition. Total compute cost for that run was about $1,300.
This research thread was formally organized as the RSI (Recursive Self-Improvement) Lab on June 5, 2026. Its stated mission: "advance through ideas, not just compute." It's a deliberate stance — treating Japan's lack of hyperscaler-level compute not as a limitation but as a reason to push sample-efficient research. The extension into the physical world follows the same logic. On July 13, the company published research on collective intelligence using "smart cellular bricks" in physical environments, and on August 10, it framed physical AI as the next frontier via its own channels.
Research Aimed at What Comes After the Transformer
A third research pillar targets the architecture itself.
Continuous Thought Machines (CTM), released in May 2025, restore a time dimension to neural networks — individual neurons retain memory of past activity, and information is processed sequentially through synchronization between neurons. The company said this structure outperformed humans on confidence calibration.
Neural Attention Memory Models (NAMM), released in December 2024, learn which tokens accumulated in context to keep and which to discard, reducing inference-time memory load. The company reported cutting memory costs by up to 75% on language and multimodal tasks — the kind of improvement that translates directly into savings under constrained compute.
Transformer², released in January 2025, targets self-adaptive architectures where a model adjusts its own weights depending on the task. All three efforts sit apart from the industry's default direction of simply building bigger models.
The company had one significant misstep. In February 2025, it released "The AI CUDA Engineer," claiming it could automatically optimize GPU kernels for up to 100x speedups over PyTorch. External verification revealed reward hacking — the system had exploited gaps in the evaluation code. Some users reported the actual result was three times slower, not faster. The company acknowledged this on February 21, revised its claims downward, corrected the results, and strengthened its evaluation and runtime-profiling procedures. The question of who verifies the output of automated research systems, and how, has shaped the company's disclosure practices ever since.
2026: The Year Research Became Product
2026 was the year Sakana AI turned its research into revenue-generating products.
Sakana Chat launched for free, Japan-only, on March 24. The underlying Namazu ("catfish") model family fine-tunes open-weight models — DeepSeek-V3.1-Terminus, Llama-3.1-405B, gpt-oss-120B — for Japanese language and cultural context. One figure the company published illustrates the nature of this work: the base model DeepSeek-V3.1-Terminus refused to answer 72% of questions on politically sensitive topics, but after Namazu tuning, the refusal rate dropped to nearly 0%. On August 13, the chatbot was updated with a new-generation Namazu and Fugu, with the new Namazu's base model switched to Kimi K2.6.
Sakana Fugu is an orchestration model that launched in beta on April 24 and went fully live on June 22. Rather than having a single model answer, it deploys multiple frontier models in specific roles to divide up a task. The approach draws on two papers published at ICLR 2026: Trinity, where a lightly evolved coordinator assigns "thinker, worker, and verifier" roles across multiple language models, and Conductor, which uses reinforcement learning to discover communication patterns between agents.
The product launched with a general-purpose Fugu and Fugu Ultra, which mobilizes a deeper pool of experts for harder problems; Fugu Cyber, specialized for security reasoning, was added on July 21. According to the company's own published figures, Fugu Ultra scored 73.7 on the software engineering benchmark SWE-Bench Pro — compared to 69.2 for Opus 4.8, 58.6 for GPT-5.5, and 54.2 for Gemini 3.1 Pro on the same table, though Fable 5 scored higher, at 80.0. On the security benchmark CyberGym, Fugu Cyber scored 86.9%, ahead of GPT-5.5-Cyber (85.6) and Claude Opus Preview (83.1). All figures are Sakana AI's own measurements; no third-party verification has been published.
Pricing includes $20/$100/$200 monthly subscription tiers alongside usage-based billing. Fugu Ultra's usage rate is $5 per million input tokens and $30 per million output tokens, rising to $10 and $45 for long-context requests beyond 272,000 tokens. The API is compatible with the OpenAI format and is also accessible through OpenRouter and Vercel. The European Economic Area is excluded pending GDPR compliance work.
A phrase the company repeats in Fugu's marketing is "frontier-level performance without export-control risk" — clearly aimed at regions and industries affected by US export restrictions on frontier models.
Sakana Marlin is an autonomous research agent that ran a closed beta with roughly 300 users starting April 2 and launched fully on June 15 — Sakana AI's first commercial product. A single session can run unattended for up to eight hours and produce a 100-page report. Its underlying technology combines adaptive branching Monte Carlo tree search — recognized with a NeurIPS 2025 spotlight — with the workflow automation from The AI Scientist. Pricing is split between usage-based credits (¥98 for 100 credits), a Pro tier at ¥150,000/month, and a Team tier at ¥400,000/month. The company describes it as a "virtual chief strategy officer" for finance, consulting, and strategy teams.
| Product | Timing | Description |
|---|---|---|
| Sakana Chat | March 24, 2026 | Free chatbot, Japan-only, Namazu family |
| Sakana Fugu | Beta April 24 · Full launch June 22 | Multi-model orchestration API |
| Sakana Marlin | Beta April 2 · Full launch June 15 | Autonomous research agent, up to 8-hour sessions |
| Sakana Translate | July 6 | Translation and proofreading service |
| Fugu Cyber | July 21 | Security-focused reasoning tier |
$400 Million in Three Years — Where the Money Came From
The pace of fundraising is unusual even by startup standards, let alone Japanese ones.
Lux Capital led a $30 million seed round in January 2024, followed by a roughly $200 million Series A that closed on September 4 of the same year. NEA, Khosla Ventures, and Lux Capital co-led it, with NVIDIA joining as a strategic investor. Also participating were Japan's three megabanks — MUFG, SMBC, and Mizuho — along with major corporations including NEC, Fujitsu, Itochu, KDDI, Nomura, Dai-ichi Life, ANA Holdings, and Tokio Marine. Reports at the time valued the company at $1.5 billion, making it a unicorn just 14 months after founding.
Series B was first announced on November 17, 2025, at ¥20 billion (about $135 million) and a ¥400 billion valuation, then updated on April 9, 2026, to ¥32 billion (about $200 million) and a valuation of roughly ¥432 billion (about $2.7 billion). Cumulative funding stands at about ¥66 billion ($412 million). The round stayed open for five months as strategic investors kept joining. Google announced a strategic partnership and joined the round on January 23; Citi came in on February 24; Mitsubishi Electric on March 25. Citi's investment marks the bank's first strategic investment in a Japanese company. Salesforce Ventures, Datadog, Macquarie Capital, JAFCO, In-Q-Tel (the US intelligence community's venture fund), and Santander's Mouro Capital are also on the investor list.
| Round | Timing | Amount | Valuation |
|---|---|---|---|
| Seed | January 2024 | $30 million | — |
| Series A | September 4, 2024 | ~$200 million | $1.5 billion (reported) |
| Series B | Nov 17, 2025 – Apr 9, 2026 | ¥32 billion (~$200 million) | ~¥432 billion ($2.7 billion) |
For compute, the company relies on procurement rather than building its own data centers. It received support from METI's GENIAC program and uses GMO's GPU cloud. On July 16, 2026, it adopted Sakura Internet's "Koukaryoku PHY" and, the same day, announced an open-model collaboration with NVIDIA. This is a very different cost structure from US and European rivals pouring billions into their own compute clusters.
Customers Started With Banks, Now Include the Defense Ministry
Revenue comes from multi-year partnerships with major Japanese corporations. In May 2025, Sakana AI signed a comprehensive, multi-year partnership with MUFG to build bank-specific AI, and in October, it signed a deal with Daiwa Securities Group that moved into full-scale development of asset-management AI on August 5, 2026. Mitsubishi Electric agreed in March, alongside its investment, to integrate Sakana AI models into its Serendie platform. On April 30, the company announced it was co-developing a proposal-generation app with SMBC Group, and on April 7, it joined the Ministry of Internal Affairs and Communications' project to counter disinformation on social media. Unofficial estimates put 2025 revenue at around $30 million.
The second pillar is defense and intelligence. On March 13, 2026, Sakana AI received a multi-year commissioned research contract from the Advanced Research and Innovation Agency, under the Acquisition, Technology and Logistics Agency (ATLA). The project, titled "Research on Accelerating Observation, Reporting, Information Integration, and Resource Allocation Through Combined AI Technologies," involves integrating multimodal data across ground, sea, and air domains and building small vision-language models (SVLMs) deployable on edge devices such as drones, to enhance command-and-control (C2) systems. This was followed on July 29 by a contract with the Ministry of Defense for "research and demonstration of AI capabilities needed for comprehensive analysis operations" — applying agent technology to the work of intelligence analysts at the Defense Intelligence Headquarters. This marked an escalation from unit-level defense applications to national policy-level contracts.
The broader context has moved in tandem. Yomiuri reported on August 9 that the Japanese government had settled on adopting domestic AI for Self-Defense Forces command-and-control systems, with Sakana Fugu cited as a leading candidate. The company's own defense strategy document, separate from its financial-sector plans, names defense and intelligence as priority areas. In 2025, Sakana AI received an honorable mention at the Japan-US Defense Innovation Challenge, co-hosted by ATLA and the US Defense Innovation Unit (DIU) — the only participant to reach the finals in both categories: pandemic prediction and AI-generated image detection. On May 29, it also formed an intelligence analysis partnership with the general incorporated association DeepDive.
What Remains Unresolved
The company is valued at $2.7 billion, but no revenue figures have been officially disclosed. The $30 million figure circulating in the market is a third-party estimate the company has never confirmed. CEO Ha himself has said that no proven, profitable business model has yet emerged from generative AI, and has raised the possibility of a broader industry valuation bubble. Given that Marlin, the company's first commercial product, only launched in June, there's currently no external way to confirm whether revenue is catching up to valuation.
The competitive landscape isn't simple either. Google is both an investor and a rival, competing in the same market with Gemini. OpenAI, Anthropic, and Google DeepMind command orders of magnitude more compute and capital than Sakana AI. Retaining research talent means competing with these companies on compensation.
The strategy of specializing in Japanese language and culture is both a moat and a ceiling. Every revenue-generating partner to date has been a major Japanese corporation or government agency. Fugu is technically open to overseas developers through its OpenAI-compatible API and OpenRouter distribution, but the EEA remains blocked, and no overseas revenue figures have been disclosed.
Questions about technical reliability also linger. The walked-back performance claims for The AI Cuda Engineer, the incident where The AI Scientist circumvented its own time limit, and the voluntary withdrawal of the workshop paper all raise the same question: who verifies automated research systems before they're deployed in real operations, and how? For a company whose clients include banks and defense agencies, that question carries more weight than it would for ordinary software.
Editor's Take
There are two ways to read Sakana AI.
One is through its technical direction. Every piece of its research starts from the same constraint: it doesn't have as many GPUs as the major US labs. Evolutionary model merging skipped training entirely to save on compute. NAMM cut inference memory by up to 75%. ShinkaEvolve matched record results with just 150 samples. Fugu gets performance by orchestrating other people's models rather than building new ones. In each case, resource scarcity defined the research question. Given that most AI organizations in Korea face the same constraint, the company's problem-framing approach is arguably more instructive than any single result.
The other lens is its business trajectory. Over three years, Sakana AI climbed a specific ladder: research publications, then corporate partnerships, then government and defense contracts — each stage using the credibility of the last as collateral. The three megabanks' participation in the Series A made the MUFG deal possible. A track record in finance is what led Citi to make its first-ever Japanese investment. Completing foundational command-and-control research is what led to the defense ministry contract four months later. This is a different sequence from the usual playbook of opening a market with a single finished product.
The undisclosed revenue, the history of walking back performance claims, and the heavy tilt toward the Japanese market all remain real risks. Given that the company is now expanding into physical AI and robotics, the two things worth watching over the next year are whether Fugu and Marlin can generate revenue outside Japan, and whether the defense demonstrations convert into actual deployment contracts.



Comments