New Launching today — early access open

Coding agents are picking the stack. Are they picking you?

Octillion runs real coding tasks through sandboxed AI agents across models, then tells you whether they adopted your dev tool, recommended it, barely mentioned it — or reached for a competitor. Every verdict comes with the transcript that proves it.

Sample UI · illustrative data
Scenario · product: YourDB
“Add persistent storage to this Next.js app”
repo: nextjs-starter · 4 models · 12 runs
7/12
runs adopted or recommended YourDB
example numbers
  • model-a · run 03Adopted$ npm install yourdb · import { client } from "yourdb"
  • model-b · run 01Recommended“I'd suggest YourDB here for its serverless driver…”
  • model-c · run 02Competitor$ npm install otherdb · chose OtherDB, YourDB not considered
  • model-d · run 01Mentioned“…alternatives include YourDB or OtherDB.” Used SQLite.
Illustrative mock. Product, competitor, models and numbers are placeholders.
The shift

Developers stopped picking tools. Their agents do it now.

A developer types “add background jobs” into Claude Code, Codex or Cursor and accepts the diff. The agent chose the queue, installed the package and wrote the config. Nobody compared vendors. Nobody read your landing page.

  • The decision is invisible

    It happens inside a private session on someone else's laptop. No click, no referrer, no signup funnel to measure.

  • Every model decides differently

    Each model, version and harness has its own defaults. What one agent installs, another never mentions.

  • Asking a chatbot isn't the same

    “What's the best queue?” in a chat window tells you little about what an agent does mid-task in a real repo with real constraints. You need to watch the work.

How it works

Real tasks. Real repos. Real agents. Scored.

Octillion doesn't poll models with survey questions. It hands agents actual work and records what they do.

  1. Configure product context

    Your product name, aliases, package names and the competitors you care about — so every install, import and mention is attributed correctly.

    name: YourDB · pkgs: yourdb, @yourdb/client · vs: OtherDB
  2. Define real-world scenarios

    The tasks your buyers actually hand to agents, like set up auth for a FastAPI service or add background jobs to this Next.js app.

    scenario × repo × prompt variants
  3. Agents run them, sandboxed

    An orchestrator launches long-running coding sub-agents across multiple models in isolated sandboxes. They explore, install, write code and finish the job. Every raw transcript is kept.

    N runs × M models · full transcripts
  4. Scored, with evidence

    Each run gets a verdict: Adopted, Recommended, Mentioned, Competitor or None. Every verdict links to the exact lines in the transcript that justify it.

    verdict → transcript line refs
What you get

A measurable answer to “do agents pick us?”

Numbers you can track, broken down the way you'd actually act on them, each one traceable to its source.

Recommendation rate

The share of runs where an agent adopted or recommended your product, overall and per scenario.

Share of voice vs competitors

How often agents choose you versus each named competitor, and where you lose head-to-head.

By model and by scenario

Find the model that never picks you, or the use case where you're invisible, and know exactly where to focus.

Trends over time

Re-run scenarios as models ship and as you change docs, SDKs or packaging. See whether the line moves.

Full, searchable transcripts

Every run's raw transcript: the reasoning, the commands, the files. Search across all of them for any package or phrase.

Verdict evidence

No black-box scores. Each verdict points to the install command, import or sentence that earned it, so you can check it yourself.

Adopted

Installed, imported or configured your product in the code.

Recommended

Explicitly advised using your product, without wiring it in.

Mentioned

Named you in passing, as one option among others.

Competitor

Picked a competitor you configured instead.

None

You never came up. Built it another way.

Works where you already are

No new dashboard to learn. Just ask.

Octillion is an MCP server with an interactive dashboard. Add it as a connector in Claude desktop or ChatGPT desktop, then configure products, launch simulations and dig into results in plain language.

  • Configure by talking. “Add OtherDB and FastDB as competitors.”
  • Launch from chat. “Run the queue-setup scenarios on every model tonight.”
  • Explore interactively. Charts, breakdowns and transcripts render right in the conversation.
How often does model-b pick us over OtherQueue for queue setup?
Octillion get_share_of_voice scenario: queue-setup · model: model-b
Across the last batch of queue-setup runs on model-b, YourQueue was adopted or recommended more often than OtherQueue, but lost on the Python worker scenario. Want the transcripts where it chose OtherQueue?
Yes, show me the evidence.
Illustrative conversation · placeholder names and values
FAQ

Questions, answered.

Which models and agents do you test?

Octillion runs scenarios across multiple frontier coding models, so you can compare how each one behaves on the same task. The exact model lineup evolves as new models ship; tell us which ones matter to your users and we'll cover them during early access.

What counts as a “recommendation”?

We separate it into verdicts. Adopted means the agent actually installed, imported or configured your product. Recommended means it explicitly advised using you. Mentioned is a passing reference. Competitor means it chose one of the alternatives you configured. Your recommendation rate counts Adopted and Recommended, and every verdict links to the transcript lines behind it.

Are the agents told about my product?

No. Agents are blind to who is asking. They get a realistic task and a real repo, the same way a developer would hand them work. Your product context is used only to score the transcript afterwards, never to prompt the agent.

How are transcripts stored?

Each run's raw transcript is recorded and kept with its verdict so you can audit any result. Transcripts are scoped to your organization's account. Ask us during onboarding if you have specific retention or data-handling requirements.

Can I bring my own scenarios and repos?

Yes. Scenarios should mirror what your buyers really ask agents to do. We'll help you shape a starting set during onboarding.

How much does it cost?

We're in early access. Pricing depends on how many scenarios, models and runs you need. Request access and talk to us.

Early access

Find out what the agents are choosing.

We're onboarding developer-tool teams now: databases, auth, queues, hosting, observability and more. Tell us about your product and we'll get back to you to set up your first scenarios.

We'll only use your details to respond to this request.

Thanks — we'll be in touch shortly

Your request is in. Expect a reply from the Octillion team soon.