Recommendation rate
The share of runs where an agent adopted or recommended your product, overall and per scenario.
Octillion runs real coding tasks through sandboxed AI agents across models, then tells you whether they adopted your dev tool, recommended it, barely mentioned it — or reached for a competitor. Every verdict comes with the transcript that proves it.
A developer types “add background jobs” into Claude Code, Codex or Cursor and accepts the diff. The agent chose the queue, installed the package and wrote the config. Nobody compared vendors. Nobody read your landing page.
# developer, 11:42pm > add background jobs to this app agent ▸ reading package.json, src/… agent ▸ I'll use SomeQueue — it fits this stack. agent ▸ $ npm install somequeue agent ▸ created src/jobs/worker.ts agent ▸ ✓ done. 3 files changed. # was your product considered? # you'll never see this session.
It happens inside a private session on someone else's laptop. No click, no referrer, no signup funnel to measure.
Each model, version and harness has its own defaults. What one agent installs, another never mentions.
“What's the best queue?” in a chat window tells you little about what an agent does mid-task in a real repo with real constraints. You need to watch the work.
Octillion doesn't poll models with survey questions. It hands agents actual work and records what they do.
Your product name, aliases, package names and the competitors you care about — so every install, import and mention is attributed correctly.
The tasks your buyers actually hand to agents, like set up auth for a FastAPI service or add background jobs to this Next.js app.
An orchestrator launches long-running coding sub-agents across multiple models in isolated sandboxes. They explore, install, write code and finish the job. Every raw transcript is kept.
Each run gets a verdict: Adopted, Recommended, Mentioned, Competitor or None. Every verdict links to the exact lines in the transcript that justify it.
Numbers you can track, broken down the way you'd actually act on them, each one traceable to its source.
The share of runs where an agent adopted or recommended your product, overall and per scenario.
How often agents choose you versus each named competitor, and where you lose head-to-head.
Find the model that never picks you, or the use case where you're invisible, and know exactly where to focus.
Re-run scenarios as models ship and as you change docs, SDKs or packaging. See whether the line moves.
Every run's raw transcript: the reasoning, the commands, the files. Search across all of them for any package or phrase.
No black-box scores. Each verdict points to the install command, import or sentence that earned it, so you can check it yourself.
Installed, imported or configured your product in the code.
Explicitly advised using your product, without wiring it in.
Named you in passing, as one option among others.
Picked a competitor you configured instead.
You never came up. Built it another way.
Octillion is an MCP server with an interactive dashboard. Add it as a connector in Claude desktop or ChatGPT desktop, then configure products, launch simulations and dig into results in plain language.
Octillion runs scenarios across multiple frontier coding models, so you can compare how each one behaves on the same task. The exact model lineup evolves as new models ship; tell us which ones matter to your users and we'll cover them during early access.
We separate it into verdicts. Adopted means the agent actually installed, imported or configured your product. Recommended means it explicitly advised using you. Mentioned is a passing reference. Competitor means it chose one of the alternatives you configured. Your recommendation rate counts Adopted and Recommended, and every verdict links to the transcript lines behind it.
No. Agents are blind to who is asking. They get a realistic task and a real repo, the same way a developer would hand them work. Your product context is used only to score the transcript afterwards, never to prompt the agent.
Each run's raw transcript is recorded and kept with its verdict so you can audit any result. Transcripts are scoped to your organization's account. Ask us during onboarding if you have specific retention or data-handling requirements.
Yes. Scenarios should mirror what your buyers really ask agents to do. We'll help you shape a starting set during onboarding.
We're in early access. Pricing depends on how many scenarios, models and runs you need. Request access and talk to us.
We're onboarding developer-tool teams now: databases, auth, queues, hosting, observability and more. Tell us about your product and we'll get back to you to set up your first scenarios.
We'll only use your details to respond to this request.
Your request is in. Expect a reply from the Octillion team soon.