NERVA × Agentic Shopping Assistant on AWS

A governance & decision-quality
layer for agentic commerce.
For retailers who refuse to deploy
their largest competitor's objective function.

Amazon just commodified the AI-commerce stack. Differentiation moves up the stack — to brand, to trust, to decision quality. NERVA is the layer that sits between the assistant and the customer's actual interest, scoring every decision across five states: COMMIT, HOLD, WAIT, CONSULT, TOXIC.

The landscape NERVA layers onto.
Amazon ASA · Reference market
$12Bincremental, 2025
Amazon ASA · Customers reached
300Mvia Alexa for Shopping
Amazon ASA · Conversion lift
3.5×conv. vs. keyword
NERVA · Pilot window
90dpre-committed falsification
01 Sheet 01 — The Gap & The Framework Visual Summary

ASA was built to convert.
Your business is built on something else.

Three categories of decision live outside ASA's optimization function. Each is a structural gap a retailer at sub-enterprise scale cannot fill alone — and one Amazon is structurally disincentivized to close.

What ASA does well

Battle-tested conversational commerce, delivered through AWS.

Models, infra, and learnings from Amazon's own retail business — customizable to your catalog and brand voice. Real, hard to replicate, available now.

300M customers · $12B incremental annualized sales · 3.5× session conv.
What it leaves uncovered

Three decision categories ASA does not govern.

Operator decisions about the AI — thresholds, upsell aggression, confidence floors. Customer decision quality — is the shopper right 90 days later? Brand-objective drift — invisible without a governance layer.

Optimized for the session. Not for the 180-day repeat rate.
Why Amazon won't fix it

A layer that suppresses purchases is incompatible with retail incentives.

The position is structurally available only to challenger brands — and to the infrastructure that serves them. The category will be defined by whoever ships first.

First mover defines the trust narrative.

The NERVA kernel — five states, one brake, one lift score.

Field-tested · 90-day discipline · Published
Commit

Evidence sufficient. Stakes contained. Entropy low.

Proceed. The kernel doesn't slow down good decisions.

ActionShip
Hold

Evidence not yet sufficient. Specific gap defined.

Returned with the missing evidence stated. Re-score when satisfied.

ActionGather
Wait

Stakes & evidence acceptable — timing is wrong.

Queued with a defined release condition. Price cycle, season, data window.

ActionQueue
Consult

Irreversible enough to require a second opinion.

Escalated above the ops team with explicit sign-off required. Default for one-way doors.

ActionEscalate
Toxic

Don't proceed. Stakes exceed evidence — or the pattern itself is the problem.

Blocked with a written rationale. The rationale becomes a precedent.

ActionBlock
Inputs to the kernel

Three signals decide which state a decision belongs to.

Entropy
0.68
Stakes
$340 · ret.
Evidence
3 indep.

Entropy measures whether the decision-maker can articulate what they want without naming a brand. Stakes weight financial, reputational, and relational exposure. Evidence weight only counts independent sources — a single opinion restated three times still counts as one.

Governing primitives

One-Way Door brake

Any high-stakes irreversible decision escalates to CONSULT or TOXIC by default — even when evidence weight is high.

Lift score · the pilot's core output
To be measured

Hit-rate of decisions that followed the kernel vs. decisions that overrode it, measured across the 90-day window. No number is claimed before the pilot runs. Tracked openly and reported at day 90 — including, if it comes to it, a negative result. The willingness to be falsified is the product.

How the three inputs map onto signals a retailer already tracks.

Deployable alongside ASA · no new instrumentation
Entropy →

Customer uncertainty signals.

Cart hesitation, repetitive or contradictory search queries, rapid category-switching within a session. The shopper cannot yet articulate what they want without naming a brand.

Already instrumented in session logs and search telemetry.
Stakes →

Item value & purchase finality.

Order value, custom-order or non-returnable categories, sizing-dependent fit, and downstream commitment chains (subscriptions, large gifts, bundled items).

Already in catalog metadata and order-flow rules.
Evidence weight →

Independent corroboration.

Review density and independence, sizing/spec clarity, third-party validation, and presence of comparable purchase history for the same shopper.

Already in PDP, review systems, and CDP history.
02 Sheet 02 — Two Configurations Where NERVA layers in

One kernel. Two surfaces.
Different risk profiles, different stories.

Configuration A is the lowest-risk deployment — a governance layer over ASA configuration changes, operator-facing only. Configuration B is the higher-narrative deployment — a decision panel surfaced inside the customer's conversational session. The pilot recommends sequencing A first, then B.

Score every config change
before it ships to production.

Before any change to recommendation thresholds, pricing rules, conversational flow, or upsell triggers ships to ASA, the change is scored by the NERVA kernel and routed by state. COMMIT decisions flow normally. Everything else gets a guardrail — not a checkpoint.

Surface
Operator dashboard. No customer-visible UI.
Risk profile
Lowest. Reversible at every step.
Time to signal
Day 30 — first directional lift read.
Pilot window
Full 90 days. Baseline for B.
  • Pre-deploy review — every config change scored before it ships. COMMIT auto-flows.
  • CONSULT escalation — irreversible-enough changes routed above ops with named owners.
  • TOXIC block — short-term conversion wins at the cost of return rate get stopped at source.
  • Operator lift score — hit-rate when the team follows vs. overrides. Tracked openly.
asa-ops.your-brand.com/nerva/queue
⌘K
Upsell threshold · raise to 0.78
change_id 41A2 · proposed 12m ago
HOLD · evidence gap defined
Two more weeks of return-rate data needed at current threshold before raising.
Entropy
0.62/ 1.0
Stakes
High· reversible
Evidence
2indep. sources
Gap to clear: Current 30-day return-rate signal at threshold 0.71 is below 21-day stability minimum. Re-score when window completes (est. Jun 14).
03 Sheet 03 — 90-Day Pilot Designed to be falsifiable

A pilot designed to end cleanly
if it doesn't work.

Duration
90days
Configs
A → Bsequenced
Reviews
Weekly+ day-45 midpoint
Pilot timeline
Config A · operator governance Config B · customer panel
Day 30 · first lift read
Day 45 · midpoint
Day 60 · Config B live
Config A
Config B
Day 0
15
30
45
60
75
90
Day 0 → 45
01

Operator layer goes live.

Configuration A deploys against a single category or workflow. NERVA scores every ASA config change before it ships.

  • Single pilot category named
  • Decision owner empowered for CONSULT
  • Read access to config + return data
  • Weekly read-outs begin day 7
Day 45
02

Midpoint review.

Operator lift score, deployment regret rate, and CONSULT/TOXIC precision reviewed against pre-committed targets.

  • Adopt · Extend · End — decided here
  • Go/no-go on Config B rollout
  • Falsification criteria re-affirmed
  • Sample-size honesty report
Day 60 → 90
03

Customer panel goes live.

Configuration B deploys on a single category or shopper segment. Return rate at 30/60/90 days is the primary metric.

  • Single category in scope
  • Conversion tracked, not optimized
  • Customer lift score visible
  • Day-90 report with explicit recommendation

Configuration A metrics OPERATOR

measured vs. matched control period

Deployment regret rate changes rolled back within 30 days
PRIMARY
Operator lift score follow-NERVA hit-rate vs. override
CORE
Decision velocity proposal → deploy time, by state
CORE
CONSULT signal precision % materially changed by reviewer
QUALITY
TOXIC signal precision % subsequently judged correct to block
QUALITY

Configuration B metrics CUSTOMER

measured vs. concurrent control segment

Return rate · 30/60/90d NERVA-panel-visible vs. control
PRIMARY
Contribution Margin per Session (CMS) per-session net profitability after returns, refunds, and direct acquisition cost — co-primary with return rate
PRIMARY
Repeat purchase rate same-customer return within 180 days
CORE
Session conversion tracked, not optimized for — judged against CMS, not against conversion alone. A slight dip in absolute conversion is expected and intended.
GUARDRAIL
Post-purchase sentiment NPS-equivalent at 7 and 30 days
CORE
Customer lift score % of NERVA pauses that did not become purchases
QUALITY
Falsification · pre-committed
All three must hold at day 90

The pilot is judged successful only if all three of the following are true.

No retroactive metric changes. If these conditions are not met, the layer is not adopted — and the pilot ends cleanly. The willingness to be wrong is the product.

Criterion 01
01

Operator lift score is positive and statistically meaningful for the sample.

Decisions that followed NERVA outperform decisions that overrode it. Effect size honestly stated against sample size; no claim of significance the data can't support.

Threshold> 0 · significant
Criterion 02
02

Deployment regret rate drops at least 20% vs. matched control period.

Of changes shipped under NERVA, materially fewer get rolled back within 30 days than were rolled back in the equivalent prior window.

Threshold≥ −20% regret
Criterion 03
03

Config B reduces return rate with only a single-digit dip in session conversion.

The customer panel can suppress some sessions — that's intended. But total revenue net of returns must improve, not collapse. Single-digit conversion dip is the ceiling.

ThresholdReturns ↓ · Conv. dip < 10%
Logic Criterion 01 AND Criterion 02 AND Criterion 03 = Adopt

Retailer commits to

Scope, access, and the integrity of the test.

  • One single named category, segment, or workflow as pilot scope.
  • Read access to config logs, return data, and session data (Config B) under standard DPA.
  • A named decision owner empowered to act on CONSULT escalations.
  • Falsification criteria accepted before pilot kickoff — no retroactive changes.

Starpoint commits to

Deployment, transparency, and an honest day-90 report.

  • Deploy NERVA as a service alongside ASA. No replacement of underlying stack.
  • Weekly read-outs, midpoint at day 45, full report at day 90.
  • Framework held neutral — no affiliate or partnership revenue inside the pilot.
  • Explicit recommendation at day 90: adopt, extend, or end. No upselling on a failed pilot.