Early Access

A comprehension engine built on Read-Time Compute.

#3 on DeepResearch Bench1
Up to 93% lower cost than frontier deep research agents2

Deep Cove Research comprehends full sources, then turns them into research reports at the level of frontier models,3 for much less cost.2

DeepResearch Bench

#3 of 15

55.29 on the official leaderboard1

A benchmark for deep research agents. It has 100 PhD-level research tasks, written by experts in 22 fields, in Chinese and English (ICLR 2026).4

45 organizations evaluate performance based on it, including:5

  • NVIDIA
  • Google
  • Microsoft
  • Amazon
  • Salesforce
  • Alibaba
  • Baidu
  • Tencent
  • ByteDance
  • Huawei
464850525456$0$5$10$15$20$25$30NotPublishedCost per Report (USD)DeepResearch Bench ScoreAlibaba (CN)Voicepica DeepResearchCellCog (US)CellCog Max · Max TierDeep Cove Research (CA)Max ReasoningOpenAI (US)GPT-6 Astra Pro Deep Research · Max ReasoningAnthropic (US)Claude Fable 5.1 Max Research · Max ReasoningxAI (US)Grok Build 4.6 · Max ReasoningGoogle (US)Gemini 3.1 Pro Deep Research · Max ReasoningOfficial LeaderboardEvaluation Based on Official RACE Method464850525456$0$10$20$30NotPublishedCost per Report (USD)DeepResearch Bench ScoreAlibaba (CN)CellCog (US)Deep Cove Research (CA)OpenAI (US)Anthropic (US)xAI (US)Google (US)Official LeaderboardEvaluation Based on Official RACE Method
Score and Cost Sources: 3, 6

Agents spend compute looping. We spend it on comprehension.

Agentic Research

85%

of input tokens are rereading session history

Frontier research agents work in a loop: search, open a page, take notes, repeat. Before each new step, they read all their old notes again. So most of their effort goes into rereading, not into new sources.2

GPT-6 Astra Pro · Max Reasoning

Read-Time Compute

2%

of input tokens are rereading session history

Our state-of-the-art orchestrator handles the loop itself, so the model does not reread old notes. We spend the savings on research: 5X more searches,2 and an engine built to read thousands of sources.7

Deep Cove Research · Max Reasoning

Looping has a price.

On average, frontier research agents cost $15.36 per task. Rereading their old notes is billed at a discount, but still costs $3.76, almost 4X our entire cost. Meanwhile, most of our cost goes to comprehending new sources.2

Agentic Research

GPT-6 Astra Pro Deep Research

$15.36

Read-Time Compute

Deep Cove Research

$0.99

93% Lower Cost
Rereading History Planning, Searching, Writing Reading New Material

Agents search the tips. We comprehend the whole source.

Agentic Research

14

Average Searches per Report

Frontier research agents have limited memory, and the loop uses most of it. So they only have room for about 14 searches per report. For many sources, they read only the short snippet, not the complete content, so details are easy to miss or misread.2

GPT-6 Astra Pro · Max Reasoning

Read-Time Compute

68

Average Searches per Report

Our state-of-the-art context management keeps the model’s memory clear. So it has room for about 68 searches per report, 5X more than frontier agents. We comprehend the whole page and turn it into evidence cards for deep comprehension.2

Deep Cove Research · Max Reasoning

Inside the comprehension engine.

A lighter model. A higher score.

The comprehension engine works from real-time, post-training sources it comprehends at runtime, not from obsolete knowledge. That makes model size a minor factor. We run a model priced about 40X lower than GPT-6 Astra Pro and Claude Fable 5.1, and score higher.8

Model CostUSD per 1M Output TokensDeepResearchBench Score$047$50GPT-6 Astra ProClaude Fable 5.1$12Gemini 3.1 Pro$6Grok Build 4.6$1.20Deep Cove Research54.59GPT-6 Astra Pro54.04Claude Fable 5.147.83Gemini 3.1 Pro53.85Grok Build 4.655.29Deep Cove ResearchModel CostUSD per 1M TokensDeepResearchBench Score$047$50$12$6$1.2054.59GPT-6 Astra Pro54.04Claude Fable 5.147.83Gemini 3.1 Pro53.85Grok Build 4.655.29Deep Cove Research

Cost and Score Sources: 8

Replicable

Frontier agents decide every step on their own: what to search, what to open and when to stop. These decisions come from the model’s pre-training, including its biases, so the same question can give a different report each time. Our engine runs every task through one research guardrail and builds every claim from post-training sources it has read. This reduces pre-training bias and makes reports more replicable.

Resumable

Frontier agents keep all their progress in one long session with the model. If the session stops halfway, the progress is lost, and you have to run the task again from the start and pay again. Our engine splits research into stages. If a run stops, you can resume it from the last checkpoint or start a new fork from there, so you never pay for the same work twice.

Scalable

You can freely scale up your deep research plan without cost pressure. Our state-of-the-art orchestrator saves the model’s context window and dices the work into many small calls, so scaling up your plan makes the cost grow linearly. Meanwhile, frontier agents pay the looping cost on every extra step, so for them the same increase makes the cost grow non-linearly.7

Light Infrastructure

Our engine handles the research process, so a light model is enough, at about 31X lower hardware cost.9

Sources

  1. 1DeepResearch Bench official leaderboard. GPT-5.5 judge, 7 Oct 2026. Deep Cove Research 55.29, rank 3 of 15.
  2. 2Deep Cove token study. DeepResearch Bench II tasks 7, 28 and 50, Oct 2026. Per task: GPT-6 Astra Pro $15.36 billed, 2.05M input tokens (85% repeated history), 14 searches; Deep Cove Research $0.99 at public API prices, 0.86M input tokens (2% repeated), 68 searches and 61 pages downloaded. Stage costs at billed and public API prices.
  3. 3Deep Cove evaluation. Official RACE method and GPT-5.5 judge, DeepResearch Bench tasks 10, 20, 30, 70 and 90, 3 judgments each, Sep 2026. Scores use the same scale as the official leaderboard.
  4. 4DeepResearch Bench paper. ICLR 2026, arXiv 2506.11763: 100 PhD-level tasks in 22 fields, Chinese and English.
  5. 5Published DeepResearch Bench results. Papers, model cards and leaderboard entries, Oct 2026.
  6. 6Cost per report. Deep Cove Research $1.04, average of the 100 official reports at public API prices. GPT-6 Astra Pro, Claude Fable 5.1: billed usage on DeepResearch Bench II tasks 7, 28 and 50. Gemini 3.1 Pro: list price, $1 to $3. CellCog: published price, up to $25.
  7. 7Engine design. Each source is comprehended in its own model call, in parallel, so capacity grows with the number of calls. Current setting: up to 384 pages per report.
  8. 8Public list prices. Per 1M output tokens, 8 Oct 2026: our model $1.20; GPT-6 Astra Pro $50; Claude Fable 5.1 $50; Gemini 3.1 Pro $12; Grok Build 4.6 $6. Scores: notes 1 and 3.
  9. 9GPU estimates. Kimi K3 model card: 2.8T parameters, MXFP4 weights, about 1.5 TB, which needs 2 servers of 8 NVIDIA B200 (192 GB each). gpt-oss-120b model card: 117B parameters, fits a single 80 GB GPU. On-demand prices, RunPod, 8 Oct 2026: B200 $6.79 per hour, H100 SXM $3.49 per hour.