Build & business document · Internal
The EV Record
A verified intelligence system for the global EV industry. Not a newsletter — a source of record that remembers who said what, when, and what happened next.
01The thesis
Trust in public content is collapsing because volume is free and verification is expensive. The scarce thing is no longer information about the EV industry. It is knowing which of it is true, and what it contradicts.
The starting observation was correct: social platforms are flooded, AI content is everywhere, and trust has migrated into small private groups. The wrong conclusion is to build a smaller social network — a vertical feed inherits every problem of public social with none of the scale that made it tolerable. LinkedIn-for-X is a well-populated graveyard.
The right conclusion is that the retreat to private groups is a demand signal for verified, synthesised, judged information — the thing those WhatsApp groups are badly approximating by forwarding screenshots to each other.
So the product is an intelligence system, and the sequence is deliberately inverted from the social-first instinct:
- Intelligence first. Valuable on day one with zero other users. No cold-start problem.
- Audience second. The readers self-identify by logging in to use the tools.
- Network last, or never. Only if readers start asking to be connected to each other.
EV is the first vertical for three reasons that compound: it is one of the fastest-growing industries globally, it is genuinely a global conversation (an Indian founder cares intensely what CATL and BYD did last week), and MI — our first client in the sector — gives us both domain credibility and a launch audience.
An earlier framing of this called the end state "one source of truth." That is the wrong goal and a dangerous one to build toward. Real industry information stays contradictory: a company guides 50k units, an analyst says 30k, actual is 38k — and all three were honestly stated at different times.
The valuable thing, and the rare one, is a source of record. It does not adjudicate truth. It tells you who claimed what, when, with what confidence, and what happened afterwards. This distinction drives the entire data model in §06 — get it wrong and the archive is worth a fraction of what it could be.
02What it actually is
Four things, in the order a reader meets them:
1 · The Brief — free, open, no login
A daily synthesis of global EV news, published at 08:00 IST. Not a link dump. Every item carries an explicit why this matters judgement, multi-source provenance, and a flag where it contradicts something previously reported. Reach is the entire point of this tier, so nothing is gated — brief, item detail, archive and company pages are all free.
2 · Delivery to where you already are — paid
The brief is free on web and in the app. Having it brought to you is the first paid tier: email newsletter, WhatsApp, or Telegram at 08:00 every morning. Same content, delivered rather than visited.
This is a cleaner split than it first appears. The reader who bookmarks the site and checks it costs us nothing and gives us reach. The reader who wants it pushed to WhatsApp has told us it is part of their working routine — that is a materially more valuable person, and a natural place for the first charge.
3 · Post generation — paid
Any item converts into a LinkedIn post in the reader's own voice. Requires login with LinkedIn profile, company and role — simultaneously the qualification mechanism, the personalisation input, and the liability control.
4 · The Ask agent — paid
Ask any question about the industry; answered from the archive with citations and dates. "What has Ather said about export markets over the last two years?" The piece with a real moat, and only good once the archive has depth.
5 · The archive — the actual asset
Everything above is a means of accumulating a structured, entity-resolved, contradiction-tracked corpus of global EV claims. Briefs can be copied. A three-year record of who-said-what cannot be — it can only be accumulated, and the only way to have it in 2029 is to start in 2026.
The line between free and paid
One sentence governs it: reading is free, everything else is paid.
| Surface | Access | Why this side of the line |
|---|---|---|
| Daily brief — web & app | FREE | Reach is the business model. Never gated, never truncated. |
| Item detail, full judgement, sources | FREE | Withholding depth would make us the thing we are replacing. |
| Archive & company timelines | FREE | Public and indexed — this is the organic acquisition engine. |
| Email · WhatsApp · Telegram delivery | PAID | Delivery into a routine. Signals a materially more engaged reader. |
| LinkedIn post generation | PAID | Produces an asset the reader publishes. Needs identity anyway. |
| Ask agent | PAID | Real marginal cost per query, and the deepest value in the product. |
| Company tier — seats, monitoring, alerts | PAID | Team deployment. The tier the Atriqa prospect actually buys. |
Which paid capabilities sit at which level, and what each costs, is deliberately not settled in this document. Ship the free tier, watch which paid surface people reach for first and how often, then price against observed behaviour rather than against a guess made before launch.
What is settled is the architecture: entitlements are per-capability from day one, so any tier structure can be configured later without a rebuild. The build does not wait on the pricing decision.
Cadence: daily at 08:00, plus a Monday week-in-review
The archive and the brief are two different clocks, and separating them is what makes daily work.
- Ingestion runs continuously. Sources are polled through the day, every day. Claims are extracted, entities resolved, contradictions detected as they arrive. Nothing is ever lost to cadence — a story breaking Tuesday afternoon is in the record Tuesday afternoon, regardless of when it publishes.
- The brief assembles at 07:15 and publishes at 08:00 IST, covering everything since the previous edition. Readers in Europe get it before their working day; the US gets it as an overnight digest.
- Monday adds a week-in-review — the same claims re-synthesised at a longer horizon, for readers who want one read rather than five. Costs almost nothing extra because the claims are already scored.
Daily cadence fails when it forces thin synthesis — some days genuinely have four items worth publishing, not eleven, and a pipeline padding to a fixed count produces exactly the generated feel we are designing against.
So: variable length, hard significance floor, never pad. Items publish because they cleared the threshold in §07, not to fill a slot. A four-item brief that says "quiet day — three things moved and here they are" reads as edited and builds trust. An eleven-item brief with seven press releases in it destroys the same trust. The floor is never lowered to make a longer edition.
Synthesis inference multiplies by roughly seven, which remains a small number — see §14. The real cost is that the calibration period in §08 becomes more important, not less: seven editions a week means seven chances a week to publish something wrong, and errors compound faster in front of the same audience. The 8–12 week review window is non-negotiable at this cadence.
03Legal design — read before anything is built
Crediting Times of India or Google News does not make reproduction lawful. It makes it attributed reproduction. What protects an aggregator is transformation and brevity, not the byline attached. This is the single most commonly held wrong assumption in this category and it needs to be designed out, not papered over.
India's framework is narrower than the US one, and that matters because we are Indian-domiciled serving a global audience. US fair use is an open balancing test. Indian fair dealing under Section 52 requires the use to fall within an enumerated purpose — research, criticism, review, or reporting current events. We design to the stricter standard, which happens also to be the better product.
The source policy: read widely, publish synthesised
The correct model separates two things that are easy to conflate. What we read is deliberately broad — the whole public conversation about the industry. What we publish is our own synthesis of facts drawn from many sources at once. Breadth of input is a quality advantage; it is the narrowness of output that keeps us safe.
So the reading list is wide: trade press, national newspapers and magazines, company and regulatory releases, press wires, analyst notes, and public community discussion where it carries signal — Reddit, industry forums, X. Multi-source reading is precisely what makes the five publishers reported this, and here is what it collectively means item possible, and that item is both the product and the legal safe harbour.
"Publicly visible on the internet" and "free to reproduce" are different things. A Times of India article is publicly readable and still copyrighted — copyright attaches on creation, not on being behind a paywall. This does not block us, because we are not reproducing it: we are extracting the facts it reports and writing our own sentences about what several sources collectively show. Facts are not copyrightable; a publisher's expression of them is.
The practical rule this produces: read anything public, reproduce nothing. That is what the pipeline already does by design — claim extraction throws the original prose away and keeps the fact, the number, the date and the link.
What we publish
- Facts, not expression. "Tata Motors announced X on Tuesday" is a fact. Facts are not copyrightable. Our sentences describing facts are our own work.
- Synthesis across sources. An item drawing on five publishers, saying what they collectively show, is original commentary. This is the legal safe harbour and the quality moat at once.
- Source links alongside every item — so a reader who wants the original article can go read it. This is credit and utility, not a legal shield, and it costs us nothing.
- Short marked quotations where a specific wording genuinely matters.
What we never publish
- Reproduced paragraphs, even with attribution
- Close paraphrase that follows one article's structure — the tell that we read one source, not many
- Full text from behind a paywall
- An unverified claim carried because a single source said it
Access, ranked by preference
Where a source offers a clean machine channel we use it, because it is more reliable and cheaper to maintain — not only for legal reasons. Preference order: official APIs and RSS → press wires → public web fetch. The last is used where no feed exists, at polite rates, honouring robots.txt, never behind a login or paywall. Where a platform's terms prohibit automated collection, we access it through its official API instead — Reddit and X both offer one.
The one rule that does not bend: paywalled content is never ingested, by any method. It is the single clearest line in this area and there is no upside to crossing it.
The post-generation carve-out
Generated LinkedIn posts are published by the reader, under the reader's own name. If that output sits close to a source article, the reader takes the reputational and legal hit. Therefore: the post generator works only from our synthesised claim record, never from source article text. This is a hard architectural boundary, not a guideline — source text is not carried into the generation context at all.
Email and data protection
Two regimes bind us simultaneously. Under India's DPDP Act, consent must be free, specific, informed and unambiguous — bundled or pre-ticked consent is prohibited outright. Under GDPR, B2B outreach can rest on legitimate interest (Art. 6(1)(f)) with a documented Legitimate Interest Assessment — but Germany and Italy require prior consent in practice, and Germany is a top-three EV market we cannot afford to mishandle.
The list is a genuine advantage but cannot be blasted. Sequence: send from a dedicated subdomain (never the primary sending domain), warm it gradually, make first contact an invitation rather than an issue, segment EU contacts out of legitimate-interest sending where local law demands consent, and grow from opt-ins thereafter.
And a commercial point that matters more than the legal one: 5,000 imported contacts at 8% open rate is worth less than 400 who asked to be there — worse as a business, and much worse as a story to a sponsor. Judge this on engaged readers, never on list size.
04Source map
We read broadly and publish narrowly. Sources are tiered by trust, and the tier travels with every claim through the entire pipeline — determining how many corroborations it needs before it can be stated flatly, and appearing in the reader-facing provenance chips.
The tiering is what lets the reading list be wide without the output being loose. A Reddit thread and a regulatory filing can both enter the system; only one of them can ever become a stated fact on its own.
| Tier | Definition | Examples in EV | Treatment |
|---|---|---|---|
| A | Primary / issuer | Company IR releases, regulatory filings, exchange disclosures, government policy notifications | Statable alone. Highest weight. |
| B | Established trade press | Electrive, InsideEVs, Electrek, Automotive News, CnEVPost, Reuters, Bloomberg | Two sources to state flatly; one → "reported by". |
| C | Analyst / research | SNE Research, BNEF, Rho Motion, ICCT, IEA | Always attributed as estimate, never as fact. |
| D | General & national press | Times of India, Economic Times, national dailies, business magazines, regional trade titles | Two sources to state flatly. Strong for India and MEA coverage. |
| E | Community signal | Reddit, industry forums, X, LinkedIn posts by identified operators | Signal only, never a sourced fact. Surfaces topics early; must be corroborated by A–D before it can appear as a claim. |
| X | Blocked | Content farms, AI-spun aggregators, unattributed rumour feeds, paywalled content | Never ingested. |
Tier E earns its place not as evidence but as an early-warning system. Operators complain about a supplier on Reddit weeks before it reaches trade press; a plant slowdown shows up in a local forum before anyone files. Treating that as a lead to investigate — never as a fact to publish — is how the brief gets ahead of incumbents rather than trailing them.
Mechanically: Tier E claims enter the archive with confidence: rumoured, are excluded from the brief by the G6 hedge gate, and either get corroborated by a higher tier later — at which point they become publishable, dated to when we first saw the signal — or quietly expire.
Geographic coverage — the actual differentiator
Most EV newsletters are regional or single-language. Someone covers US EV well, someone covers Europe, and almost nobody covers China properly rather than translating press releases. A brief that credibly spans China, EU, US, India and MEA is a different product, not a better version of the same one. It is also precisely what AI aggregation is genuinely good at, so our cost to do it is a fraction of a newsroom's.
China coverage specifically is where the gap is widest and the value highest — CnEVPost and Chinese-language primary sources give us reach almost nobody in the Indian or Gulf market has.
The competitive picture, honestly
The field is real but beatable on a different axis. Electrive (Berlin, founded 2013) is the strongest incumbent for decision-makers. EV Universe runs weekly to roughly 7,000 readers. The EV Report does daily. Numerous regional players exist.
We will not beat them on journalism. They employ people who do this full-time and love it. Content is not a moat — do not attempt to compete there.
We win on three things they structurally cannot build: global aggregation breadth (they are regional), the action layer (no newsletter turns an item into your post, or tells you which companies in the story to talk to), and the queryable archive (nobody is storing claims as structured, contradiction-tracked records).
05The archive — design principle
Everything else in this document is downstream of one decision: the archive stores claims, not facts.
A naive design collapses information into current state — one row per company with a units_2026 field that gets overwritten as new numbers arrive. This destroys history, and history is the entire product. The moment you overwrite March's guidance with July's revision, you have lost the only interesting thing: that they revised.
So: every assertion is stored as a timestamped, attributed, immutable claim. Nothing is ever overwritten. Claims link to other claims through typed relationships — corroborates, contradicts, supersedes, updates. Current state is derived by querying the claim graph, never stored.
If verification treated consistency with the archive as truth, a wrong figure entering in month two and being repeated would, by month six, cause the system to confidently flag the correction as the anomaly — because the correction contradicts everything on file.
Therefore: contradiction triggers a flag, never a suppression. An outlier is surfaced for judgement, never silently dropped. "This contradicts what was reported in March" is frequently the most valuable item in a brief.
Entity resolution is the underinvested piece
"Tata Motors", "Tata Motors Ltd", "TML", and "टाटा मोटर्स" must resolve to one entity or the archive is worthless for querying. This is unglamorous work that determines whether §11 is possible at all. Entities carry aliases, tickers, parent/subsidiary relations, and sector tags, and every claim links to resolved entity IDs — never to raw strings.
06Schema
Postgres. The core is five tables; the shape below is the load-bearing part.
Three details carry disproportionate weight. asserted_at separate from ingested_at — a claim made in March that we ingest in July must sort by March. confidence as an enum on every claim, so hedge language in §09 is derived mechanically rather than decided by a model. And numeric_value extracted with its unit and period, never generated — the single most common way an AI pipeline embarrasses itself is inventing a plausible number.
07Judgement — the significance rubric
An aggregator without judgement covers everything equally: a press release about a dealership opening gets the same weight as a strategic reversal. The fix is not a better prompt. It is an explicit rubric every candidate claim is scored against, with the scores stored and auditable.
| Dimension | Question | Weight |
|---|---|---|
| Decision impact | Does this change what a supplier, OEM or investor does next week? | 30 |
| Money moving | Capital committed, capacity built, contract signed — with a number attached? | 20 |
| First / reversal | First of its kind, or a reversal of a previously stated position? | 20 |
| Breadth | Does it affect a segment, or one company? | 15 |
| Contradiction | Does it conflict with something in the archive? | 15 |
Everything is scored. The top items publish; the rest are discarded from the brief but retained in the archive. The discipline of discarding is what makes a brief feel edited rather than generated — and because discarded claims still enter the archive, nothing is lost for §11.
Three required editorial moves
Each of these exists because its absence is a specific, recognisable failure of automated briefs.
- The dissent slot. One item per brief must argue that something widely covered does not matter. Forcing a position is what separates judgement from summary.
- The archive callback. Where an item connects to prior claims, the brief says so explicitly with dates. This is the visible proof that the system remembers.
- Lead variation. The opening item alternates in form across weeks — a number, a contradiction, a question, a reversal. Identical structure every week is why readers stop opening after a month.
Anti-repetition
The pipeline reads its own archive before writing. Reporting a partnership as new that was covered three weeks ago destroys credibility faster with industry readers than almost any other error — and they are exactly the readers we want. This is a database problem, not an AI problem, and it is entirely solvable.
08Verification — seven gates
Every claim passes these in order. A failure at any gate either downgrades confidence or drops the claim; nothing skips ahead.
Source admissibility
Tier X blocked outright. Paywalled content rejected at fetch, by any access method. Feeds and APIs preferred; public fetch permitted at polite rates where no feed exists.
Entity resolution
Every named organisation resolves to a canonical entity, or the claim is held for review rather than published against a guessed match.
Numeric extraction
Numbers are extracted with unit and period and traced to the exact source sentence. Any figure that cannot be traced is dropped — never estimated, never carried forward.
Corroboration
Tier A stands alone. Tiers B and D need two independent sources to be stated flatly; one source becomes "reported by X". Tier C is always framed as estimate. Tier E can never satisfy corroboration — community signal raises a topic for investigation but cannot itself become a published fact.
Archive cross-reference
Matched against prior claims on the same entity and metric. Contradictions and numeric deltas generate claim_link rows and raise significance rather than suppressing the claim.
Hedge enforcement
Language is derived mechanically from confidence. A rumoured claim cannot render in declarative voice. This is a code path, not a prompt instruction.
Voice pass
Press-release register stripped. Corporate boilerplate, superlatives and announcement-speak removed before publication.
For the first 8–12 weeks, one person reviews the assembled brief before it sends. Ten minutes. Not writing — calibration: learning what the rubric gets wrong and feeding it back.
This is not a human in the loop forever. It is a human in the loop until the loop is trustworthy. Going fully hands-off from issue one means the errors are discovered by readers — on a list of people we intend to sell to, in an industry where a single confident error about a competitor's numbers is remembered for years.
Corrections are handled visibly: a wrong item publicly corrected costs far less than one quietly deleted, and the correction itself becomes a claim in the archive.
09Brief assembly — worked example
What a single item looks like when it comes out of the pipeline. This is the product's core mechanic rendered as it would actually appear.
CATL holds 39.9% of global battery share while BYD slips to 14.4%
H1 2026 installation data puts CATL at 39.9% of global EV battery deployments, roughly flat against the 40.2% recorded through May. BYD fell to 14.4% from 16.7% across the comparable 2025 period — the wider of the two moves, and the one worth attention.
Archive callback: BYD held 16.7% in Jan–Nov 2025 (CLM-2013, reported 7 Jan 2026). The 2.3-point fall is a sustained trend across four consecutive readings, not a single-quarter artefact.
Three things are doing work here that a summariser cannot do. The why it matters is a position, not a restatement. The archive callback proves memory and is the item's most valuable sentence. And the provenance strip shows tier, count and gate status — the reader can audit us, which is the entire trust proposition.
10Interface
Every screen — the free brief, email, the paywall moment, sign-up, post composer, the Ask agent and its refusal state, company timeline pages and the mobile app — is laid out screen by screen in the companion document.
The design direction is an instrument panel, not a magazine. Dark ground, monospace for anything that is data or provenance, serif for the synthesis prose — because the serif signals editorial judgement, which is exactly what we are selling. Sodium-amber as the single accent, used only for judgement and action.
Web — the daily brief
Overnight from China: CATL's H1 share holds while BYD slips again — the fourth reading in the same direction. Europe filed charging numbers that miss their own Q2 target. Quiet otherwise, so this is a short one.
CATL holds 39.9% of global battery share while BYD slips to 14.4%
Combined 54.3%. BYD's fall is the fourth consecutive reading in the same direction.
EU public charging additions fall 12% against Q2 target
Reported by two national operators; the Commission has not commented. Treat as directional until confirmed.
Why yesterday's solid-state announcement does not matter yet
Pilot-line, no volume commitment, no date. Covered everywhere this morning, covered wrong.
App — the Ask agent
Three distinct positions, and they have shifted.
In March 2025 BYD confirmed the Hungary plant with a stated 2026 start1. By January 2026 that had moved to "phased through 2027" without the change being announced directly — it surfaced in a supplier briefing2. Most recently, July 2026, a second European site was described as "under evaluation", with no country named3.
The pattern worth noting: each statement has been less specific than the one before it. No date has been formally withdrawn, but none has been reconfirmed either.
Two rules govern the agent's interface. Every factual sentence carries a citation to a claim ID with its date and confidence — an uncited assertion is a bug. And where the archive is thin, it says so and declines, rather than generating a plausible answer. A confident wrong answer here costs more than every right answer gains.
Mobile and email
Email is the primary distribution surface and is designed first: the summary paragraph and item headlines must work as plain text with no images. The app is a reader plus the two paid tools — post generation and Ask — not a feed. There is no social layer, by design.
11The Ask agent
Retrieval over the claim store, not over article text. This distinction is what makes it defensible and what keeps it legal.
A query resolves entities first, then retrieves candidate claims by hybrid search — pgvector embedding similarity combined with structured filters on entity, date range, geography and claim type. Retrieved claims are ordered chronologically, and the model's job is to narrate the record, not to answer from knowledge. Its context contains claims with their IDs, dates, confidence levels and sources; it has no access to source article text.
Hard constraints
- Every factual sentence cites a claim ID
- Confidence language inherits from the claim's
confidencefield - Contradictions in the retrieved set are surfaced, never resolved silently
- Below a retrieval-coverage threshold the agent declines and says what it lacks
- No answering from model knowledge — if it is not in the archive, we do not know it
This is the smallest component to build and the most dependent on what precedes it. Built against a thin corpus it is worse than not shipping it at all, which is why it is sequenced last.
12Money
The honest framing: this is a lead-generation and credibility engine for Atriqa, wearing the clothes of a media product. Media monetisation alone is brutal. As a front door to services revenue, the economics invert entirely.
Which changes the success metric. Not subscribers. How many of the target accounts read us daily. Three hundred of the right people beats thirty thousand of the wrong ones.
"I'm from Atriqa" gets ignored. "I write the EV Record you read every morning" gets a meeting. Every reader is a warm prospect who has watched us demonstrate competence sixty times before we pitch. Engagement data — who opens, who clicks what, who forwards — makes this a ranked list, not a blast list. Two retainers a year pays for the whole thing many times over.
Reading is free everywhere. Everything else carries a charge: delivery into email, WhatsApp or Telegram; post generation; the Ask agent; and a company tier with seats, competitor monitoring and alerts.
The delivery tier is the interesting one commercially — it is the lowest-friction paid thing in the product, because a reader asking for the brief on WhatsApp at 8am has already told us it belongs in their routine. That is a much easier first charge than a tool subscription, and it identifies the most engaged readers for us. Exact tiers and prices are deferred deliberately — entitlements are per-capability from day one so any structure can be configured later without a rebuild.
Not CPM. Sold as named sponsor of the brief the industry's decision-makers read. B2B newsletter benchmarks run $25–150 CPM with specialised niches commanding $2,000+ per placement — but at our scale one sponsor at a real number beats twenty at small ones, because one sponsor is a sales conversation rather than an ad-ops business. Natural candidates: charging infrastructure, battery suppliers, EV-focused funds, and event organisers — the last being home turf. MI is the obvious first conversation, once we have leads to show them.
The audience and the archive together make an industry event straightforward to convene and to sell. Deliberately out of scope for now, but it is where this most plausibly becomes a business rather than a channel.
No paid subscription on the brief. A paywall kills reach, and reach is the entire point when the real product is who reads it. Also note the deliberate tension: attribution means every item links out. That is correct here — we capture relationship value, not attention-minutes — but it means this must never be measured on time-on-site.
13Build plan
Four layers, sequenced so each is useful before the next exists.
Ingestion & archive
Source registry with tiers, RSS/API/wire pollers, deduplication, entity resolution, claim extraction, the claim store. Accrues value from day one whether or not anyone reads a brief.
Node · Postgres + pgvector · scheduled workers
Editorial engine
Significance scoring, multi-source clustering, contradiction detection against archive, synthesis with mechanical hedge enforcement, voice pass. Output: the brief.
LLM pipeline · rubric as code · humanizer pass
Distribution, identity & entitlements
Free web and app brief. Paid delivery channels — email on a warmed subdomain, WhatsApp Business API, Telegram bot — each fanning out the same assembled edition at 08:00. Auth with LinkedIn/company/role capture, per-person engagement tracking, post generation from claim records.
Per-capability entitlements from day one, so tier structure and pricing are configuration rather than code.
Next.js · auth · ESP · WhatsApp Business API · Telegram Bot API · entitlement service
The Ask agent
Hybrid retrieval over claims, citation-enforced generation, coverage-threshold refusal. Small to build, entirely dependent on corpus depth beneath it.
pgvector hybrid search · constrained generation
Email is the simplest and needs only the domain warming already described in §03.
Telegram is nearly as simple: a bot, users opt in by starting a chat, no per-message cost, no template approval. Ship it first of the two chat channels.
WhatsApp is the one with real constraints. It runs through the WhatsApp Business API, requires a verified business account, uses pre-approved message templates for anything outside a 24-hour user-initiated window, and charges per conversation. A daily 08:00 push is exactly the template-and-charge case. It is worth doing — WhatsApp is where this audience actually reads in India and the Gulf — but it needs approval lead time and a per-message cost line, so it cannot be treated as "email but different".
Sequencing
Lock the source list and tiers. Design and review the claim schema. These are the two decisions that are expensive to reverse — everything else is recoverable, but a wrong claims model either loses the history or forces a rebuild.
Ships: source registry, schema, one real issue produced end to end
Ingestion live across all tiers and geographies. Entity resolution working. Claims accumulating daily. Brief assembled on manual trigger.
Ships: archive filling, brief producible on demand
Web and app brief live, free. Auth and profile capture. Daily 08:00 cadence begins with the calibration review in place. Paid delivery channels (email, WhatsApp, Telegram) and post generation for early users.
Ships: free daily brief, paid delivery + composer, readers identified
Calibration converges and review moves to spot-check. Contradiction detection becomes genuinely useful as the archive thickens. Paid tiers tested against observed usage.
Ships: reliable unattended pipeline, first revenue signals
Ask launches on a corpus with real depth. This is the point at which the product becomes hard to copy.
Ships: queryable industry record
14What it costs to run
Less than intuition suggests. Ingestion and storage are cheap at this volume. The meaningful variable is LLM inference — and it scales with sources processed, not readers, which is the favourable direction: a thousand new readers cost nothing to serve the brief to.
| Component | Scales with | Character |
|---|---|---|
| Ingestion & storage | Sources × time | Negligible. Text is small. |
| Claim extraction | Items/day | The main running cost. Bounded and predictable — a few hundred items daily. |
| Brief synthesis | Per issue × 7/wk | Small per edition; ~7× weekly at daily cadence. Still modest. |
| Post generation | Per use | Small, and directly attributable to a paying user. |
| Delivery — email & Telegram | Recipients × 7/wk | Email near-zero at this scale. Telegram free. |
| Delivery — WhatsApp | Conversations | The one real per-message cost. Priced per conversation; must be covered by the delivery tier. |
| Ask agent | Queries | The only cost that scales with engagement. Small at year-one volumes. |
| Infrastructure | Flat | Sits on existing server capacity. |
It is the 8–12 weeks of calibration attention — watching what the rubric gets wrong and tuning it. That is the input that determines whether this reads as authoritative or as generic. It cannot be bought, only spent. Budget it explicitly or the product will be mediocre for reasons that have nothing to do with the technology.
15Risks, ranked honestly
| Risk | Severity | Mitigation |
|---|---|---|
| Publishing something wrong Rumour as fact, misread number, LOI reported as signed deal | Highest | Seven gates (§08). Numbers extracted never generated. Calibration period. Visible corrections. |
| Reads as AI slop Even-weighted, no memory, no position, padded to length | High | Significance rubric, dissent slot, archive callbacks, lead variation, humanizer pass. |
| Entity resolution failure Aliases fragment the archive | High | Alias tables, held-for-review on ambiguity, never publish against a guessed match. |
| Copyright exposure | Medium | Read broadly, reproduce nothing. Synthesis over paraphrase, multi-source by construction, no paywalled content, post generator firewalled from source text. |
| Community signal treated as fact A forum rumour reaching the brief as a claim | Medium | Tier E cannot satisfy G4 corroboration and is blocked from the brief by the G6 hedge gate. |
| Email reputation damage From mishandling the launch list | Medium | Dedicated subdomain, gradual warming, invitation-first, EU segmentation. |
| Thin editions padded to length Quiet days filled with press releases | High | Hard significance floor, variable length, never pad. A 4-item brief is a valid brief. |
| Nobody cares | Low–Med | Phase 0 produces one real issue before meaningful spend. That is the honest test. |
16Open questions
Four things this document deliberately does not settle, because settling them requires information we do not yet have.
MI's role
Currently: first client, source of the launch database, reason EV was chosen. The open question is whether they become a launch partner or first sponsor. The sequencing instinct is right — build first, show leads, then have the conversation from a position of demonstrated value rather than asking them to fund a concept.
Exact composition of the launch list
Size, fields, geography and how contacts were sourced. This determines whether the launch advantage is real or theoretical, and it changes the outreach design — particularly the EU segmentation described in §03.
Pricing
Deliberately unresolved. Ship the free tier, watch what people actually use and how often, then price against observed behaviour. The company-tier instinct is stronger than the individual tier but should be tested, not assumed.
Second vertical
The architecture is vertical-agnostic — sources, entities and rubric weights are configuration. But multi-vertical early is a trap: each one is a separate content pipeline, a separate credibility problem, and a separate sales motion. Prove EV to real revenue before templatising. The platform story is attractive on a slide and lethal in execution.
Prepared for Atriqa. Competitive and legal research current as of this date.
Next action: Phase 0 — lock source registry, review claim schema, produce one real issue.