Skip to content
research & development

We built the dataset because nobody had measured it

We publish what we measure. The work below exists because we needed the answers for client builds and found no trustworthy source — so we built the dataset ourselves.

programmes

Four lines of work, each one already in a product

01Live · 1,140 stores

The Indian D2C cohort

A classified benchmark of Indian D2C storefronts, segmented by region, category, sub-segment and AOV bracket. It exists because every available conversion and trust benchmark is drawn from US and EU stores, which makes them useless for a brand selling COD into Tier-3.

Outputs

  • 60+ scored rules across trust, UI/UX, regional fit, SEO and AI Search
  • Twelve India-specific signals — COD, UPI, GSTIN, pincode, WhatsApp, festival timing
  • Cohort-relative scoring rather than global averages

Powers every D2CIQ audit.

02Open methodology · first round pending

Agent evaluation

A reproducible test suite for agents answering customers on commerce stores. Six probe families, weighted scoring, published transcripts. Pre-registered so the tests cannot be shaped around a result.

See it →

Outputs

  • Six probe families including privacy leakage and escalation judgement
  • Worst-of-three scoring to surface non-determinism
  • Published transcripts, including our own failures

Published as the D2C Agent Benchmark.

03Live · scoring in production

AI Search visibility (GEO)

Why answer engines cite some storefronts and ignore others. We score structured data coverage, server rendering, entity completeness and FAQ markup, then measure what actually changes citation behaviour.

Outputs

  • GEO scoring dimension inside the D2CIQ audit
  • JSON-LD, FAQPage, SSR and entity-coverage checks
  • Applied to this site — see /llms.txt and our own schema

Shipped as the GEO dimension in D2CIQ.

04Early · one vertical shipped

Constraint modelling for product advice

Most product finders ask about specifications. Customers do not think in specifications — they think in constraints. We model the constraint space of a category, then map it to the catalogue.

See it →

Outputs

  • Five-question advisor mapping kitchen constraints to chimney models
  • Questions posed in customer language, bilingual English and Hindi
  • Rendered as a theme section so merchandisers edit it, not engineers

Shipped for Cravia.

open questions

Things we do not know yet

Published because a research page that only lists answers is a brochure.

  1. 01Does deflection rate correlate with satisfaction at all, or does it mostly measure customers giving up?
  2. 02How much of an agent's regional failure is the model, and how much is missing grounding data?
  3. 03Can escalation judgement be evaluated without a human in the loop, or is it irreducibly a judgement call?
  4. 04What is the actual half-life of a golden eval set before catalogue drift makes it misleading?

Want the cohort data applied to your store?

The audit scores you against the 1,140-store benchmark in about sixty seconds. No signup.

Talk to an engineer