Exa Agent Ultra in 2026: Pricing, API, Benchmarks, and When to Use It

By

· Published

· Updated

5 views

·

, ,
Exa Agent Ultra editorial cover showing a research agent coordinating many sources into structured results

Exa has launched Agent Ultra, a new highest-effort mode for its Agent API aimed at research jobs where completeness matters more than speed or a fixed per-request price. Ultra can coordinate multiple agents across thousands of sources, run for up to three hours, and spend up to a configurable budget while returning cited, structured research results.

The release matters because Exa is moving beyond “search API” territory into a direct competitor to frontier-model deep-research systems. For developers, analysts and AI-agent builders, the key question is not whether Ultra can search the web — Exa already did that — but when a long-running, metered research agent is worth using instead of ordinary Search, Deep Search, fixed-effort Agent runs, Perplexity, Claude or GPT-based research workflows.

This guide covers what changed, how Agent Ultra works, current pricing and limits, the API call, benchmark claims, practical tradeoffs, and when Ultra is actually the right tool.

What changed with Exa Agent Ultra?

Exa Agent already supported several fixed effort levels and an automatic mode. Ultra adds a new top tier designed for exhaustive research rather than predictable cost or low latency.

Before Agent Ultra With Agent Ultra
Agent effort could be fixed at minimal, low, medium, high or xhigh, or left on auto. Developers can set effort: "ultra" for Exa’s highest-compute research mode.
Fixed modes emphasized predictable cost per request. Ultra is metered by actual Agent usage and can run much longer.
Typical Agent research was optimized around bounded effort levels. Ultra is intended for large list building, deep multi-source research and hard-to-verify criteria.
Research jobs were generally expected to complete within ordinary agent-runtime windows. Ultra typically takes around 30 minutes on complex work and can run for as long as three hours.
Budget control was mostly about choosing an effort tier. Ultra adds explicit cost and duration caps so teams can bound a long-running job.

Ultra is not a separate model. It is an orchestration mode for Exa Agent: the system breaks a job into subtasks, assigns research work across multiple agents, uses Exa’s web index and tools, and spends more compute when the task requires it.

How to call Agent Ultra

The API change is simple. On an Exa Agent run, set:

{
  "query": "Find all companies building browser automation tools in the United States.",
  "effort": "ultra"
}

Exa’s documentation says the rest of the Agent request remains compatible with existing capabilities such as structured output schemas, input data and streaming.

For long-running jobs, developers should not rely on a normal short polling window. Exa’s own examples use a three-hour polling timeout, and the API also supports streaming events and background-style workflows.

Agent Ultra pricing and budgets

Agent Ultra is not priced like Exa’s fixed effort tiers. It is metered at Exa’s standard Agent usage rates.

The default maximum spend for an Ultra run is $20, but a run that finishes early costs less. Exa also lets developers set their own budget controls:

Control Allowed range Default
budget.maxCostDollars $1 to $100 $20
budget.maxDurationSeconds 300 to 10,800 seconds (5 minutes to 3 hours) No explicit duration cap

For example:

{
  "query": "Build a complete market map of European battery-recycling companies with founders, funding and customers.",
  "effort": "ultra",
  "budget": {
    "maxCostDollars": 10,
    "maxDurationSeconds": 1800
  }
}

As a run approaches either limit, Exa says the Agent stops starting new work and returns what it has found.

This is a meaningful difference from ordinary fixed modes, where the base request price is known in advance. Exa’s public pricing page currently lists fixed Agent modes from $0.012 per request for minimal through $1.00 per request for xhigh. Ultra is the escalation path for tasks that justify variable spend.

Current Exa Agent pricing

Effort mode Base price / billing model Best fit
Minimal $0.012/request Lightweight lookups
Low $0.025/request Simple factual research
Medium $0.10/request Routine research
High $0.50/request Harder research with more citations
X-high $1.00/request High-value work where completeness matters
Ultra Metered usage; default cap $20/run Exhaustive research, large lists and difficult verification

Exa’s Agent usage can also include separately billed search calls and optional enrichment services. Teams should verify the current pricing page before putting Ultra into a high-volume production workflow.

How long does Agent Ultra take?

Exa says a complex Ultra task typically completes in about 30 minutes, while very challenging work can take up to three hours.

That makes Ultra a poor fit for interactive chat where a user expects an answer in seconds. It is better suited to asynchronous jobs: diligence reports, large market maps, dataset construction, literature reviews and enrichment pipelines that can be queued and collected later.

If latency is the priority, Exa’s ordinary Search and Deep Search endpoints remain a better architectural fit. Ultra should be thought of as a research job, not a faster search request.

What Agent Ultra is designed to do

Exa highlights three broad classes of work:

  • Model and AI teams: collect papers, repositories, benchmarks and training-data candidates while verifying difficult criteria.
  • Financial and research teams: build market maps, perform due diligence, investigate KYC evidence and monitor portfolio-company events.
  • Go-to-market teams: build large account lists, enrich rows with custom judgment-based fields and attach citations for each result.

The common thread is recall. Ultra is most useful when missing relevant entities is expensive. If the task only needs five good examples, ordinary Agent modes are usually cheaper and faster.

Agent Ultra vs Exa Search, Deep Search and Agent

Exa product What it does best Typical speed / cost profile
Search API Retrieve relevant web results for an application or agent. Low latency; priced per search request.
Deep Search Multi-step web-grounded answers with structured outputs and citations. Seconds-to-tens-of-seconds class; priced per 1,000 requests.
Agent fixed effort Research tasks with bounded, predictable per-run spend. $0.012–$1.00 per fixed-effort request.
Agent Ultra Exhaustive list building and deep research where completeness dominates latency. Metered; typically ~30 minutes, up to 3 hours; default $20 cap.

The decision is therefore less “which endpoint is smartest?” and more “how much recall, latency and budget does this job justify?”

Exa’s benchmark claims — and how to read them

Exa reports that Agent Ultra leads the systems it compared on WANDR, DeepSearchQA, WideSearch and an internal Find-All Company benchmark. The company’s published results include:

Benchmark Exa Agent Ultra Selected comparison
WANDR soft recall 81.4% Opus 5.5: 72.3%
DeepSearchQA 93.9% GPT-6 Astra: 85.3%
WideSearch 58.9% Perplexity Agent: 56.0%

These figures are useful, but they should be read as vendor-reported benchmarks, not as an independent guarantee that Ultra will outperform every competing product on your workload.

Exa explains that its WANDR evaluation uses the upstream grader logic but changes the contents tool, transport logic and judge model, and that some competitor results were published by those vendors while other runs were executed by Exa. That methodology is more transparent than a bare benchmark claim, but teams should still test on their own acceptance criteria before replacing an existing research stack.

For comparison context, AI-XBlog also tracks Claude Opus 5.5, OpenAI GPT-6 Sol and Luna, and Perplexity pricing and API costs.

Why the budget controls matter

The biggest operational risk with an exhaustive research agent is not only hallucination. It is uncontrolled effort.

If a system can keep branching into more searches, more sources and more verification tasks, “be thorough” can turn into an unpredictable bill or a long-running job that does not materially improve the answer.

Ultra’s explicit cost and duration limits create a useful production boundary. A team can define, for example:

  • routine market research: fixed high or xhigh;
  • important diligence: Ultra with a $5–$10 cap;
  • rare exhaustive investigations: Ultra with a larger budget and multi-hour timeout.

This tiered approach is usually better than sending every research request to the maximum setting.

What happens when Ultra hits a limit?

Ultra exposes clear stop reasons. Exa documents outcomes such as:

  • schema_satisfied — the research task completed;
  • budget_reached — the run hit the cost cap;
  • time_limit_reached — the run reached the duration cap;
  • stopped — the developer stopped it manually;
  • error or cancelled — the run did not complete normally.

A run can also be stopped early while preserving what it has already found, with billing limited to usage up to the stop.

OpenAI Responses API compatibility

Exa also documents Agent Ultra through an OpenAI-compatible Responses API path. When using that interface, developers can set reasoning.effort: "ultra" with streaming or background execution enabled.

This is strategically important because it reduces migration friction for applications already built around the Responses API pattern. The research backend can be swapped or tested without redesigning the entire application around a proprietary request shape.

When should you use Agent Ultra?

Use Ultra when missing results is expensive

Examples include regulatory research, diligence, company-universe construction, evidence-backed dataset creation and investigations where a partial list creates downstream risk.

Use a fixed effort mode when you need predictable unit economics

If you run thousands of similar research jobs, a known $0.10, $0.50 or $1.00 request cost is easier to forecast than a metered run.

Use Search or Deep Search when the user is waiting

Ultra’s 30-minute typical runtime makes it unsuitable for most synchronous assistants, customer-facing search boxes or interactive support experiences.

Use your own benchmark before trusting vendor comparisons

Take 20–100 real tasks, define what counts as a complete and correct answer, and compare recall, citation quality, runtime and total cost against the system you already use.

Practical implications for AI-agent builders

1. Web research is becoming its own compute tier

AI developers increasingly have to choose not just a model, but a research budget. Agent Ultra makes that tradeoff explicit: spend more time and compute when the task requires completeness.

2. Research agents need asynchronous architecture

A three-hour maximum runtime means production systems should use background jobs, status polling, streaming events, retries and durable result storage rather than request-response assumptions.

3. Cost controls should be part of the prompt-routing layer

The application should decide which tasks qualify for Ultra. Letting every user request select maximum effort directly is a simple way to create unpredictable spending.

4. Citations are necessary but not sufficient

For high-stakes research, teams still need source-quality rules, date checks, duplicate handling and human review. An exhaustive agent can retrieve more evidence, but it cannot make weak sources authoritative.

5. Ultra competes with frontier models on workflow economics, not only intelligence

The value proposition is that specialized search infrastructure plus orchestration can outperform a general-purpose frontier model on wide-and-deep retrieval tasks. Whether that advantage holds depends heavily on the workload.

Bottom line

Exa Agent Ultra is a meaningful expansion of the Exa Agent API because it introduces an explicit “run to exhaustion” tier for web research.

It is available now with effort: "ultra", typically takes around 30 minutes on complex research, can run for up to three hours, and is metered with a default $20 maximum spend per run. Developers can set their own cost and time caps, stop runs early, use structured outputs and integrate through Exa’s native Agent API or an OpenAI-compatible Responses API path.

The strongest use case is not ordinary Q&A. It is research where completeness is worth paying for: market maps, due diligence, list building, enrichment and multi-source verification.

For routine jobs, Exa’s fixed effort modes remain easier to budget. Ultra should be the escalation tier — not the default.

Sources

Last updated: September 26, 2026. Pricing, limits and preview behavior can change; verify Exa’s current documentation before deploying a production workflow.

AI-XBlog Weekly Brief

Keep up with AI that actually works

Join the AI-XBlog Weekly Brief for major AI updates, practical workflows, useful tools, and editor’s picks. No daily noise.

Double opt-in. Unsubscribe anytime. See our Privacy Policy.

Reader discussion

Join the discussion

Have you tried this tool or workflow? Share your experience, corrections, or questions. Useful reader feedback may help us improve this article.

All comments are reviewed before publication. Your email address will not be published. Promotional links and low-value spam are removed.

Add a comment

Comments are moderated to keep the discussion useful and trustworthy.

About the author

AI-XBlog Editorial Team researches and maintains practical coverage of AI tools, automation, agents and applied artificial intelligence. We prioritize primary sources, clear evidence and useful real-world guidance.

Editorial Policy · Review Methodology · Corrections Policy