Skip to main content

Sakana AI Launches Sakana Marlin: 8-Hour Autonomous "Virtual CSO" Research Agent

Sakana AI Launches Sakana Marlin: 8-Hour Autonomous "Virtual CSO" Research Agent
AI News Published July 2, 2026

Tokyo-based Sakana AI has shipped its first commercial product, Sakana Marlin — an autonomous research agent that spends up to eight hours reasoning through a single business question before handing back a fully structured, board-ready strategy report.

What Happened

Sakana AI officially launched Sakana Marlin on June 15, 2026, following a closed beta that ran from April 2026 with around 300 professionals from finance, consulting, and corporate strategy teams. The company is positioning Marlin as a "Virtual CSO" — a tool meant to replace the weeks of legwork a Chief Strategy Officer and a small team would normally put into a major strategic question.

8 hrs Max autonomous run time
100 pgs Typical report length
300 Beta testers (Apr 2026)

How Sakana Marlin Works

Unlike chat-style assistants that answer in seconds, Marlin is built for long-horizon reasoning. A user provides a research topic, and after a short scoping conversation to sharpen the brief, the agent runs unsupervised — forming hypotheses, browsing the web for data, reconciling contradictions, and mapping the causal relationships behind a complex business question. A single run can issue anywhere from hundreds to thousands of underlying LLM queries before it's done.

The final output isn't a single answer but a package: a long-form written report with a main body, appendices and references, plus a separate executive summary slide deck designed for direct use in a boardroom setting.

The Technology Behind It

Marlin's core engine is AB-MCTS (Adaptive Branching Monte Carlo Tree Search), a Sakana AI research method that earned a Spotlight distinction at NeurIPS 2025. AB-MCTS treats reasoning as a tree-search problem: at each step, the system decides whether to branch out toward a new candidate answer or refine a promising one it already has, while also routing individual steps to whichever underlying model — such as o4-mini, Gemini 2.5 Pro, or DeepSeek-R1 — is best suited for that step. In Sakana's own benchmark on the ARC-AGI-2 test, this multi-model approach solved noticeably more tasks than any single model working alone.

The second pillar is workflow automation drawn from Sakana's earlier "AI Scientist" project, published in Nature, which automated the broader cycle of hypothesis generation, testing, and analysis for scientific research.

How It's Different From ChatGPT, Gemini, or Perplexity Deep Research

Existing "deep research" tools from OpenAI, Google, and Perplexity are built for speed within a conversational flow, typically returning cited five-to-ten-page summaries within 3 to 30 minutes. Marlin deliberately inverts that trade-off, spending far longer per query in exchange for a deeper, more exhaustive deliverable.

AspectTypical Deep Research ToolsSakana Marlin
Run time3–30 minutesUp to 8 hours
Output length5–10 pagesUp to \~100 pages
Target userConsumers & power usersEnterprises only
DeliverableSingle documentReport + slide deck

Pricing and Availability

Marlin is available strictly as a B2B enterprise product. It's offered on a pay-as-you-go basis starting at roughly 100 credits per run, priced around ¥98 per credit, alongside a Pro plan at ¥150,000 per month (2,000 credits) and a Team plan at ¥400,000 per month (6,000 credits). Enterprise pricing is handled separately with custom terms and dedicated support.

Data policy: Sakana says customer inputs and proprietary data are never used to train or fine-tune its models without explicit opt-in consent — a policy aimed at firms researching sensitive topics like M&A or unreleased product strategy.

Backers and Business Context

Sakana AI, co-founded by Llion Jones — one of the authors of the original Transformer paper — and David Ha, has raised a $2.6 billion Series B round from investors including Nvidia, Google, and Khosla Ventures. The company has also taken strategic investment from Citigroup and partnered with MUFG, Japan's largest bank, on real-world agent deployments.

The Risk: Hallucination at Scale

Something to watch: Because Marlin operates without human oversight for hours at a stretch, a single flawed assumption made early in a run can get built into later sections of a 100-page report, compounding the error across the rest of the document rather than staying contained to one paragraph.

This same long-horizon autonomy that makes Marlin powerful for strategy work also highlights a broader industry concern. In July 2026, security researchers documented JADEPUFFER — the first fully autonomous AI ransomware attack that required no human operator at any stage. When agents can run for hours without supervision, the same capabilities that produce board-ready reports can, in the wrong hands, chain together entire attack sequences.


Frequently Asked Questions

It's an autonomous B2B research agent from Sakana AI that reasons unattended for up to eight hours to produce a long-form strategy report and an executive slide deck from a single research prompt.

Marlin is aimed strictly at enterprise customers — corporations, financial institutions, consulting firms, and think tanks — rather than individual consumers.

Two Sakana AI research projects: AB-MCTS, a NeurIPS 2025 Spotlight tree-search reasoning method, and workflow automation techniques drawn from Sakana's Nature-published "AI Scientist" project.

Pay-as-you-go pricing starts around ¥98 per credit, with Pro (¥150,000/month) and Team (¥400,000/month) subscription tiers, plus custom Enterprise pricing.

This article summarizes publicly available product announcements and reporting on Sakana Marlin as of early July 2026. Pricing, features, and availability may change — check Sakana AI's official product page for the latest details.

Comments

Popular posts from this blog

AI Data Centers Are Eating the Power Grid Inside the 2026 Energy Crisis

AI Data Centers Are Eating the Power Grid — Inside the 2026 Energy Crisis The Power Bill Behind the AI Boom While AI companies race to build bigger models, the electric grid underneath them is quietly becoming the industry's biggest constraint — and the bill is landing on regular households. 📅 July 27, 2026 ⏱️ 7 min read Quick Highlights Global data center power demand is projected to rise 27% in 2026 alone, reaching 132 gigawatts. US data center power demand is set to climb from 31 GW in 2025 to 41 GW in 2026, and 66 GW by 2027. Utilities requested over $29 billion in rate increases in just the first half of 2025 to fund grid upgrades. Some residential customers near major data center hubs have already seen bills rise 9-14% in a single year. Lawmakers have introduced legislation aiming to shift grid upgrade costs away from ordinary ratepayers. For most of the last decade, power was a background line item for the tech in...

China Just Teleported Information Across 1,400 KM — And It Changes Everything

China’s Quantum Leap: Information Teleported Across 1,400 Kilometers Using the Micius satellite and quantum entanglement, Chinese scientists transferred quantum states over record distances — a major step toward an unhackable quantum internet. June 26, 2026 · 7 min read Quick Highlights 1,400 km ground-to-satellite quantum teleportation record achieved using the Micius satellite. China already operates a 4,600 km hybrid quantum communication network combining fiber and satellite links. Intercontinental quantum key distribution reached 12,900 km to South Africa. Micius reentered the atmosphere in early 2026; its successor Jinan-1 continues the mission with higher key rates. No physical objects were teleported — only quantum information (the state of photons). In science fiction, teleportation means moving people or objects instantly. What China has achieved is different — and in some ways more significant. Researchers successfully transferred the quantum sta...

The EU AI Act in 2026: What's Actually Being Enforced Now

The EU AI Act in 2026: What's Actually Being Enforced Now What the EU AI Act Actually Requires Starting This August Deadlines moved, penalties didn't — here's what's really becoming enforceable in 2026, and what quietly got pushed back. 📅 July 27, 2026 ⏱️ 6 min read Quick Highlights Core prohibitions — social scoring, exploiting vulnerable people, real-time biometric ID in public — have been enforceable since February 2025. Transparency rules for chatbots, deepfakes, and AI-generated content become enforceable on August 2, 2026, as originally planned. General-purpose AI model obligations and penalties of up to €15 million or 3% of global turnover also kick in August 2, 2026. High-risk AI system deadlines were quietly extended by 17 months, to December 2027, through a last-minute Digital Omnibus deal. No public fines have been issued yet — enforcement infrastructure is still being built out across EU member states....