A living document. Add findings as we learn them, with a date and a source. When something turns out to be wrong, correct it in place and say so — a struck-through wrong belief is more useful than a silently deleted one, because it stops us rediscovering it.
Last updated: 3 Oct 2026
Published as the Research library: outputs/admin-site/research-library.html (generated — run bash outputs/build-research-library.sh after editing). Served on the admin origin only, never the customer app. Future research belongs in this file or in outputs/docs/*-research.md, then rebuild; do not leave it scattered. Competitor research is maintained in Google Drive (see §3).
Overview
Last updated not dated
1. What the product is, and where it actually wins
The product: ingredient/supplier management, recipe costing, and derived PPDS allergen labels for small UK baking businesses.
Table stakes (everyone has these — they don't win anything): PPDS labels with bolded allergens · allergen matrix · recipe costing · audit trail · sub-recipes.
Our actual moat — two things, neither of which competitors advertise:
Per-bake supplier selection. The label proves which pack actually went in the mixer, not which pack the recipe assumes. Swap to an alternate supplier on the day and the label — and its allergens — change with it. Everyone else derives labels from the recipe; we derive from the batch. This is the defensible idea.
Second-opinion allergen scanner + freshness tracking (v0.11.0, v0.13.0). Nobody in the comparisons below sells "we check our own allergen detection and tell you when it's stale".
Strategic read (17 Jul 2026): deepen the moat, don't chase the feature list. Stock control and order management are where competitors are strong and we'd be late; correctness and trust is where we're already ahead and where a labelling product lives or dies.
Market
Last updated 31 Jul 2026
2. Market
Who: UK small food businesses — not bakers specifically. PPDS is not a bakery rule: the FSA publishes sector guidance for butchers as well as bakers, and meal-prep caterers, farm shops, delis and cafes packing sandwiches, market stalls and street food, chocolatiers, and jam, preserve and sauce makers are all in scope on identical terms. Home bakers and cottage bakeries are one vertical among several, not the market. Solo or family-run; a laptop and a phone; no IT support. Price-sensitive.
(Corrected 29 Jul 2026. This paragraph led with "UK home bakers" and read as a baker-first market, which understated the addressable set and contradicted §11 and stella-apps-business-research.md §1a — both of which establish that widening beyond bakers is copy and positioning, not build. Now logged as D-24.)
The buying trigger is almost certainly compliance fear (Natasha's Law), not efficiency. That suggests the pitch is "prove you did it right", not "save time".
Biggest adoption barrier — evidenced, not assumed: entering ~70 ingredient declarations by hand. Our own instance has 70 sourcings; nobody will type that to trial an app. This is why supplier product capture and parsing is so important.
⚠️ The arithmetic here is stale, and the update strengthens the conclusion — see §17 (24 Aug 2026). The instance now has 181 sourcings, of which 91 still carry the (enter declaration) placeholder and 157 have never been confirmed against a pack — a month after #249 shipped a working one-tap pack reader. The barrier is not just the first 70; per-item capture does not clear a backlog at all.
Open question: market size unknown. Not yet researched.Sized 31 Jul 2026 — see §12 and gtm-channels-research.md §1. UK register ≈ 540k rated / 605k total; TAM ≈ 150–240k PPDS-obligated businesses (£27–43m ARR); realistic 3-year SOM 1,000–5,000 subscribers (£180k–£900k ARR). Two load-bearing numbers remain modelled rather than known: the PPDS fraction and the active home-based stock (~60–120k, modelled).
Full report: outputs/docs/stella-apps-business-research.md (six sections: market expansion, must-have UK SaaS features, competitive landscape, suggested pricing, HMRC/business setup, banking and accounting). It carries its own source list. Headlines worth having here:
Competitive anchor confirmed and sharpened. FoodCore is still £19 Essentials / £55 Core, no per-label fee, no per-user fee for small teams, 7-day trial with no card. (⚠️ Superseded 10 Aug 2026: FoodCore repriced to three tiers, £25/£40/£65 inc. VAT — the anchor is now £25–65/mo and our under-price opening widened. See §3 and the master competitor sheet.) Nutritics is reported at roughly £80-200/mo and is not aimed at our buyer. US home-baker tools sit at $20-49 (Bakesy $20/$40 with a free tier, BakeOnyx $29, CakeBoss $49/$99/$179); Craftybase runs $49-349. (Corrected 20 Aug 2026, live vendor pages, during the website comparison re-verification: Bakesy dropped its free tier — now Standard $9.99 / Premium $17.99 with a 30-day trial; and Craftybase rebranded to Stocksmith (stocksmith.io) — same team, software and $49–349 prices, craftybase.com redirects there. FoodCore re-verified unchanged at £25/£40/£65 inc. VAT: matrix and food safety still Core-only, costing/orders/stock still Growth, AI checks still metered 20/50/100, 7-day trial no card.)
The sub-£10 slot is empty.Bake Diary closed in May 2025. It was Ireland-based, served mostly UK and Irish bakers at about EUR 6.95/mo, and was the cheapest credible option. That is both our clearest opening and a warning: the market has just watched a cheap single-founder tool disappear, so credibility signals (status page, export promise, backup policy) are commercial features, not hygiene.
Recommended pricing: £9 Starter / £19 Standard / £39 Pro, annual at 10x monthly, 30-day full-feature trial with no card, and a read-only Archive state rather than locking data at the end. The label engine stays complete in every tier; the paywall runs along orders, accounts, receipts, stock and users. Fixed cost base is roughly £50-70/mo, so break-even is about 8 Starter or 4 Standard customers. Stripe's £0.20 fixed fee is 2.2% of a £9 charge, so annual billing nearly halves fee drag at the low tiers.
Expansion is cheapest sideways, not abroad. PPDS is not a bakery rule: meal-prep caterers, farm shops, delis, market stalls, butchers and preserve makers are in scope on identical terms, so that expansion is copy and positioning, not build. Then Ireland (same FIC base, but PPDS food is treated as non-prepacked under S.I. 489/2014 — the FSAI cites this "as updated by S.I. 656/2024"; §29a records that 656 changed enforcement, not the declaration — so a different label template), then AU/NZ PEAL (21 allergens, plain-English, prescribed format, mandatory since 25 Feb 2024), then Canada (11 priority allergens; the real cost is bilingual labels), then the US last (9 allergens, simplest label, hardest market). Candle/wax-melt CLP labelling is the strongest non-food adjacency: same shape exactly, a mandatory label derived from supplier documentation, with the same insurance-invalidation fear as the buying trigger. (⚠️ Order corrected in place 6 Sep 2026, #825 — outputs/docs/decisions/825-choose-markets-and-order.md, register entry D-29, DECIDED — Dave confirmed on #825: the UK's own jurisdiction gap comes first. The food-safety module seeds SFBB, an FSA England-and-Wales pack, and ProvenBot is hard-wired to food.gov.uk — a Scottish (Food Standards Scotland / CookSafe) or Northern Irish (Safe Catering) business is misdefaulted today, and D-01 says "UK", not "England and Wales". Filed as #1349. The corrected order: UK adjacent verticals (now, copy only) → Scotland and NI (#1349) → Ireland (first non-UK, #826 → #827, post-GA only) → AU/NZ → Canada → USA last, re-tested and confirmed on market grounds. Candles/CLP is removed from the market order — a product line on the same engine, not a market, decided on its own issue. Two things the table above does not show: Ireland's regime for PPDS food is lighter than Natasha's Law, so the compliance buying trigger is weaker there and Ireland is first for the abstraction and the Bake Diary story, not for revenue (~8% of the UK by population); and B2C digital sales into Ireland carry Irish VAT from the first euro, so "ready for Ireland" includes choosing non-Union OSS or a merchant of record before #827 starts — see the VAT bullet two entries below.)
To make any of that cheap, three things must become data not code: the declarable allergen set, the label template, and the locale layer. (Corrected 27 Jul: the original version of this bullet said #48 single tenancy gated everything. It does not. #48 shipped in v0.14.0 - see the correction in §6. The real launch gates are billing, staging (#137), the legal stack and support tooling.)
Market size is still not closed.Closed to a defensible range on 31 Jul 2026 — see §12. Two corrections to this bullet's figures: the register is ~540k rated / ~605k total (the 430,000+ was the older rated-only count), and the 37%-domestic figure is confirmed and dated (37% of 92,540 registrations Mar 2020–Mar 2022, FSA campaign). The FSA still publishes no home-based stock total; §12 models it at 60–120k active.
UK SaaS obligations that actually apply to us: ICO Tier 1 fee £52/yr (£47 by DD); UK GDPR plus the DUAA (in force 5 Feb 2026, complaint-handling duties from 19 Jun 2026) means a DPA, a sub-processor list and a published complaints route; VAT threshold £90,000, frozen for 2026-27, but no threshold at all for B2C digital sales outside the UK, so stay UK-only or use a merchant of record; the DMCC subscription regime is consumer-facing and delayed to spring 2027, but build to it anyway. And say plainly in the Terms that the food business operator owns their label.
Stella Apps setup: sole trader is right for now; register for Self Assessment by 5 October after the first trading tax year; Class 2 is gone, Class 4 is 6% / 2%; cash basis is HMRC's default from 2024/25; MTD for Income Tax reaches £20,000 qualifying income in April 2028, assessed on 2026-27, so keep digital records from day one.
✅ Banking: Mettle — ACCOUNT OPENED 29 Jul 2026, along with sole trader registration at HMRC. Chosen over Tide on one specific ground - it includes the full FreeAgent platform free, which is HMRC-recognised for MTD and removes the accounting line item entirely. Tide is perfectly workable (its 20p transfer fee barely bites a business with a handful of movements a month) and Starling is the strongest full-bank alternative; hold a second account somewhere else as insurance against a frozen one.
Full write-up
From outputs/docs/stella-apps-business-research.md
Stella Apps / ProvenBatch - Commercial Research Report
Prepared 27 Jul 2026. Companion to outputs/RESEARCH.md (which stays the living record); this is the standing commercial reference for launching ProvenBatch as a paid product under Stella Apps.
Scope. Six questions: where the toolset travels with least rework, what any UK SaaS must have, who we are competing with, what to charge, how to set Stella Apps up with HMRC, and how to bank and book-keep it.
How to read the evidence. Every factual claim below is sourced in the Sources list. Where a figure could not be verified it is marked [unverified] and is not used to support a recommendation. Prices and thresholds move: re-check anything commercial before acting on it.
Not advice. This is research, not regulated tax, legal or financial advice. Confirm tax and compliance points with HMRC guidance or an accountant before filing anything.
1. Market Expansion Opportunities
The product is not really "bakery software". It is a chain: supplier declaration → ingredient → recipe → derived, provenance-backed label. Anything that shares that chain is a candidate market. Two axes: adjacent verticals (same law, so almost no rebuild) and geographies (different law, so the rebuild cost is the allergen taxonomy and label format).
1a. Adjacent UK verticals - effectively zero regulatory rebuild
Natasha's Law / PPDS is not a bakery rule. The FSA scope covers any business packing food on the premises it is sold from, and the FSA publishes sector-specific PPDS guidance for butchers as well as bakers. Businesses in scope include bakeries and cake businesses, market stalls, cafes and delis packing sandwiches, salads and ready meals, meal-prep and catering businesses packing for collection or delivery, home bakers selling online or at markets, and farm shops and food halls packing their own produce.
Vertical
Why it fits today
What actually needs changing
Meal prep / batch-cook caterers
Identical PPDS duty; recipes, sub-recipes and per-batch supplier swaps are the same shape
Wording ("bake" → "batch"/"cook"), portion sizing
Farm shops and food halls
Pack own produce on site; often many small lines, few staff
High line count, daily changes, high allergen exposure
Faster daily label reprint flow
Market stalls and street food
Pack before market day; price-sensitive; mobile-first
Already served by the offline capture work (M1-M3)
Butchers
FSA publishes PPDS guidance specifically for them; heavy compound-ingredient use (sausage mixes, marinades)
Weight/pack terminology; "use by" handling
Chocolatiers, jam and preserve makers, sauce makers
Same compound-ingredient and allergen problem, smaller recipe count
QUID percentage and net quantity (already logged as gaps)
Read: this is the cheapest expansion available and it does not need a single regulatory change. It is a positioning and copy exercise, not a build. The one caveat is that FoodCore already names "bakeries, meal prep businesses, home producers, caterers and market traders", so this ground is contested rather than empty.
Market size, honestly. The Food Hygiene Rating Scheme covers more than 430,000 businesses across England, Wales and Northern Ireland. The FSA reports that 37% of new registrations since March 2020 are run from domestic kitchens at private addresses, and separately warns that many home-based sellers have never registered at all. That supports "tens of thousands of home-based food businesses" as a plausible order of magnitude, but the FSA does not publish a home-based total, so the precise addressable number remains unverified (this is RESEARCH.md §7 open question 1, still open). Note the double edge: unregistered sellers enlarge the latent market but are, by definition, the least compliance-motivated buyers.
1b. Geographic expansion, ranked by rebuild cost
Market
Regulatory basis
Rebuild cost
Notes
Republic of Ireland
Same EU Reg. 1169/2011 (FIC), same 14 allergens
Low
Ireland treats food prepacked for direct sale as non-prepacked, with national rules (S.I. 489/2014, updated by S.I. 656/2024) requiring allergen information in writing, rather than the UK's full ingredient list. So: a second label template plus EUR currency, not a new engine. Bake Diary, the incumbent serving mostly UK and Irish bakers, closed in May 2025, leaving this market unusually open
EU (wider)
Same FIC base; national rules vary for non-prepacked
Medium
Genuinely prepacked goods need the full ingredient list we already produce. Cost is language, currency, and a per-country rule table
Australia / New Zealand
Plain English Allergen Labelling (PEAL), mandatory since 25 Feb 2024
Medium
21 declarable allergens (adds molluscs, lupin), individual tree nuts named, gluten cereals named individually, plain-English terms in bold in a prescribed format and location. Recency is the opportunity: PEAL forced everyone to re-label. Cost is a wider allergen taxonomy plus a format engine
Canada
11 priority allergens plus gluten sources and added sulphites
Medium-high
Declared in the ingredient list or a "Contains" statement. The real cost is bilingual English/French labels, not the allergen logic
USA
FALCPA, 9 major allergens since sesame was added 1 Jan 2023; "Contains" statement after the ingredient list
Low on labelling, high on everything else
The label is simpler than ours. The difficulty is 50 different state cottage-food regimes, plus the most crowded competitive field (CakeBoss, Bakesy, BakeOnyx all price in USD). Enter last, if at all
1c. Adjacent non-food markets (same engine, different rulebook)
Candle and wax melt makers (CLP labelling). Structurally the closest analogue to what we already do: a mandatory hazard label derived from supplier documentation (a Safety Data Sheet and a CLP sheet for the exact fragrance percentage used), required on anything containing fragrance or essential oil, at any scale, including market stalls and Etsy. Crucially the buying trigger is identical to ours: product liability insurance commonly requires correct labelling, so a wrong label can invalidate cover. Rebuild: swap the allergen engine for a pictogram / hazard-statement engine; keep supplier declarations, derivation and provenance untouched. This is the strongest non-food candidate.
Soap and cosmetics (UK Cosmetic Products Regulation). Requires a UK Responsible Person, a Cosmetic Product Safety Report signed by a qualified assessor, and an INCI ingredient list on the packaging. We could generate the INCI list and RP block, but the gating artefact (the CPSR) is a human assessment we cannot supply, so the software is a smaller part of the customer's problem. Weaker fit than candles.
1d. What to build once, to make all of the above cheap
None of this is affordable while jurisdiction rules are baked into code. Before any expansion:
Fix single tenancy first.Corrected 27 Jul 2026, same day as first issue. The original version of this report repeated a stale RESEARCH.md §6 claim that the app is single-tenant and unsellable until #48 lands. #48 shipped in v0.14.0: tenant tables, flat RLS policies, self-serve sign-up, and isolation proven against a real throwaway tenant. There is no architectural gate on selling. The actual launch gates are billing, a separate staging environment (#137), the legal stack and support tooling. Nothing else in this report depended on the wrong claim, but §3b and §5 have been corrected to match.
Allergen list as data, not code - a jurisdiction profile carrying the declarable set (14 UK / 21 AU-NZ / 11 CA / 9 US), the emphasis rule, and the naming rule (for example PEAL's "name each tree nut individually").
Label template as data - flat ingredient list vs "Contains" statement vs written allergen notice, driven by the same profile.
Locale layer - currency, units, date format, and language strings extracted from the single HTML file.
A jurisdiction is then a config row plus a template, which is the difference between a two-week market entry and a three-month one.
Recommended order: UK adjacent verticals (now, no build) → Ireland (small build, vacated by Bake Diary) → AU/NZ → candles as a separate product → Canada → USA.
2. Must-Have SaaS Features and UK Considerations
2a. Product and platform must-haves
Grouped by whether they gate launch.
Gating (cannot sell without):
Multi-tenant data isolation with per-tenant RLS, proven by tests. Done in v0.14.0 (#48).
Subscription billing with plan gating, trials, card updates and dunning (failed-payment retries).
Data export in an open format, both for UK GDPR portability and because "you can leave with your data" is a trust argument in a compliance product.
Backups with a tested restore, not just backups.
A support route with a stated response expectation.
Near-gating (needed within weeks of the first paying customer):
Onboarding import. Our own evidence is that entering roughly 70 ingredient declarations by hand is the single biggest adoption barrier. Barcode capture (#27) and receipt scanning are therefore not features, they are the funnel.
Users and roles. If a baker's staff verify allergens, roles (#38) become a compliance question, not a convenience.
A label archive: an immutable record of exactly what was printed, when, and from which supplier batch. This is our differentiator turned into an artefact the customer can hand to an EHO.
Vendor support access that is time-boxed and logged (RESEARCH.md §6a). Decide before customer #1 asks, not during their first problem.
Operational:
Error monitoring, uptime monitoring and a public status page.
Transactional email with SPF, DKIM and DMARC configured (compliance email that lands in spam is worse than no email).
Cookieless product analytics, so no cookie banner is needed (see PECR below).
A documented incident and breach process (72-hour clock).
2b. UK regulatory and operational notes
Area
What applies
Practical action
VAT
Registration threshold is £90,000 of taxable turnover in any rolling 12 months, frozen for 2026-27
Below it, charge no VAT. That is a genuine price advantage over VAT-registered rivals: the headline price is the price. Monitor the rolling 12-month figure monthly
VAT on digital services
For B2C digital services the place of supply is where the customer is, and there is no threshold for selling into other countries
Restrict sales to the UK at launch, or use a Merchant of Record. Many home bakers are not VAT-registered and behave like consumers, so do not assume B2B reverse charge saves you
ICO data protection fee
Tier 1 is £52 a year, or £47 by Direct Debit, for micro-organisations (up to 10 staff or £632,000 turnover)
Register and pay before processing live customer data
UK GDPR / DUAA 2026
The Data Use and Access Act came into force 5 February 2026, with complaint-handling duties taking practical effect from 19 June 2026. There is no small-business exemption
Publish a privacy notice, keep a record of processing, offer a customer-facing DPA (you are a processor of their order and customer data), publish a sub-processor list (Supabase, Netlify, Stripe, email provider— D-12 carries the current list; Netlify removed 29 Jul 2026), document DSAR and breach handling, and publish a complaints route
PECR / cookies
Consent needed for non-essential cookies and analytics
Choose cookieless analytics and avoid the banner entirely
Subscription law (DMCC Act 2024)
The new subscription regime is now expected spring 2027 and applies to trader-to-consumer contracts. It will require clear pre-contract information, renewal reminders on a durable medium, easy exit and cooling-off refunds
Your customers are micro-businesses, and the line between "consumer" and "trader" is thin. Build to the spirit now: pre-contract clarity, renewal reminders, one-click cancel. Cheap to build in, expensive to retrofit
Liability
This is compliance software. The food business operator, not the software, is legally responsible for the label
Say so explicitly in the Terms. Do not market "compliant labels"; market "labels derived from your data, with an audit trail". Carry professional indemnity insurance
Accessibility
WCAG 2.2 AA is not statutory for private-sector SaaS, but the Equality Act 2010 duty to make reasonable adjustments applies
Treat AA as the internal bar; it is also cheaper than retrofitting
Contract stack
-
Terms of Service, Acceptable Use, Privacy Policy, DPA, sub-processor list, and either an SLA or an explicit statement that there is none
2c. Recommended service providers
Need
Recommendation
Cost (verified July 2026)
Why
Card payments
Stripe
1.5% + £0.20 UK standard cards; 1.9% + £0.20 premium/commercial; 2.5% + £0.20 EEA; 3.25% + £0.20 international; £20 per chargeback win or lose; refunded fees are not returned
Default for UK SaaS, no setup or monthly fee
Subscriptions
Stripe Billing
0.7% of recurring billing volume (pay as you go)
Handles dunning, proration, invoices. Avoid the £450/month plan at this scale
Tax automation
Stripe Tax only if selling outside the UK
0.5%
Unnecessary while UK-only and under the VAT threshold
International alternative
Paddle (Merchant of Record)
~5% + $0.50, includes global VAT registration, collection and remittance
Worth it only if you sell B2C outside the UK. Paddle becomes the seller of record and takes the tax liability
Hosting
Netlify (already in use)Cloudflare Workers
Free tier now credit-based (300 credits/month); Pro $19 per member per month Static-asset requests are free and unlimited on both plans; deploys are not metered
⚠️ Corrected 5 Aug 2026 — D-20 / #251 moved everything to Cloudflare Workers and Netlify is gone in fact (both sites deleted, account closed, 3 Aug 2026). The credit-based Netlify pricing is exactly why. "No change needed" has not been true since 28 Jul
Database, auth, functions
Supabase (already in use)
Pro $25 per project per month, including 8 GB database, 100k MAU, and a $10 compute credit
Free tier pauses projects after 7 days idle, so Pro is mandatory the moment a customer exists
Transactional email
Resend (3,000 emails/month free) or Postmark (from $15/month, 10,000 emails)
as noted
Resend's free tier covers launch; Postmark has the stronger deliverability reputation
Error tracking
Sentry (free tier)
£0 at this scale
Analytics
Plausible or PostHog, cookieless configuration
low / free tier
Avoids a cookie banner
Status page
Instatus or BetterStack (free tiers exist)
£0 to low
Insurance
Superscript (tech-focused, from around £10/month) or Hiscox (professional indemnity quotes from £8/month)
~£100-150/year
Essential for compliance software
Accounting
FreeAgent, free with a Mettle or NatWest account
£0
See §6
Regulator
ICO Tier 1
£52/year (£47 by DD)
Estimated fixed cost base at launch: roughly £50 to £70 per month (Supabase Pro plus insurance, ICO fee amortised, domain, with hosting and email on free tiers). That number drives §4.
3. Competitive Landscape
(⚠️ FoodCore's entry here is superseded — they repriced to £25/£40/£65 inc. VAT; the current analysis is the master competitor sheet.)
All pricing verified July 2026. Currency is as the vendor lists it; USD prices are not converted, because the exchange rate is not the point, the price perception is.
3a. Who is actually in the market
UK, direct:
FoodCore - the closest comparison and the price anchor. Two plans: Essentials from £19/month (recipes, allergens, Natasha's Law labels) and Core from £55/month (adds costing, shopping lists, orders, food safety logs, business insights, team management). No per-label charges, no per-user fees for small teams, no setup fee, no minimum contract, 7-day full-feature trial with no card. Explicitly targets bakeries, meal prep, home producers, caterers and market traders. They also run a substantial content and SEO operation, which is a competitive fact in its own right.
Planglow (LabelLogic Live) - established UK label supplier with software attached. Different business model: the software pulls through label and consumable sales. Pricing not published in a form we could verify [unverified].
Nutritics - nutrition analysis and compliance for dietitians, manufacturers, hospitals and larger caterers. Reported at roughly £80-200/month, tiered by users and products, with complex onboarding and no PPDS label printing aimed at small producers. Not a competitor for our buyer; it is who our buyer is told they should use and cannot afford.
Home-baker tools, mostly US-priced:
CakeBoss - $49/month up to 50 orders, $99/month up to 200, $179/month unlimited.
Bakesy - free up to 5 orders/month, $20/month Starter, $40/month Growth (unlimited orders).
BakeOnyx - $29/month, positioned at 10 to 40 custom orders a month, mobile-first.
Baking It - cake-business tooling (tin and serving charts, portions, 3D design, orders, quotes), reported at over 40,000 users.
Craftybase - the maker-economy COGS and inventory tool: $49 Studio, $99 Indie, $199 Business, $349 Growth per month, with about two months free on annual billing, and a 14-day trial. Priced by order lines, not by recipes.
The vacancy:Bake Diary closed in May 2025. It was an Ireland-based recipe-costing app serving mostly UK and Irish bakers at about EUR 6.95 a month, and it was by some distance the cheapest credible option. Its closure did two things: it emptied the sub-£10 tier, and it gave the whole market a reason to distrust cheap single-founder tools. Both matter to us.
3b. Feature comparison
Rows in bold are where we currently differ. "Ours" reflects what is live at v0.68.0.
Capability
Ours
FoodCore
CakeBoss
Bakesy
Craftybase
Nutritics
PPDS / Natasha's Law labels, allergens emphasised
Yes
Yes
No
No
No
Partial (manufacturer labels)
Allergen matrix and export
Yes
Yes
No
No
No
Yes
Recipe costing
Yes
Core tier
Yes
Yes
Yes
Partial
Nested sub-recipes, flattened to leaves on the label
Yes
Yes
No
No
Partial
Yes
Per-batch supplier selection reflected on the label
Yes
No
No
No
No
No
Second-opinion allergen scan plus staleness/verification tracking
Yes
No
No
No
No
No
Barcode capture / Open Food Facts import
Yes
No (not advertised)
No
No
No
No
Receipt scanning from a photo
Yes (Claude vision, with on-device Tesseract as fallback)
No
No
No
No
No
Offline capture in the kitchen (PWA, outbox)
Yes
Not advertised
No
No
No
No
Order management
Yes (v0.20.0), with prices (v0.22.0)
Core tier
Yes
Yes
Yes
No
Stock control
No (#33)
Core tier
Partial
Partial
Yes
No
Cash-basis accounts and monthly P&L
Yes (v0.22.0)
Business insights only
Partial
No
COGS focus
No
Multi-tenant with self-serve sign-up
Yes (v0.14.0)
Yes
Yes
Yes
Yes
Yes
Food safety / HACCP logs
No (#36)
Core tier
No
No
No
No
Published UK price under £15/month
n/a
No
No
No
No
No
3c. Pricing analysis
The UK anchor is £19 to £55 a month (FoodCore). Nutritics sets a ceiling reference at £80-200.
The US home-baker tools cluster at $20 to $49, with CakeBoss going to $179 for volume.
Nothing credible now sits under £10 a month in the UK. Bake Diary occupied that slot at about EUR 6.95 and it is gone.
Our buyer is a cake-shed or home baker whose business may generate "several hundred pounds per week". £55 a month is roughly a whole day's takings for some of them; £19 is arguable; under £10 is an easy yes. Price sensitivity here is real, not assumed.
Charging per label is a live model in this market (Planglow's model depends on consumables). FoodCore explicitly does not charge per label, so "no per-label fee" is table stakes, not a differentiator.
3d. Where we can actually differentiate
Batch-level provenance. Everyone else derives the label from the recipe. We derive it from the batch that was made. That is the only claim in this document that no competitor advertises, and it is the one to build the brand on. Turn it into an artefact: a printable, timestamped record of what went in.
We check our own work. The second-opinion allergen scan plus freshness and verification tracking is unique in the comparison set, and it directly answers the failure mode our own data proved is the dangerous one: silent under-declaration.
We remove the data entry. Barcode capture and photo receipt scanning attack the evidenced adoption barrier (roughly 70 declarations to type). No competitor in the set advertises either. This is a funnel advantage, which is worth more than a feature advantage.
On-device by default. Receipt photos are never uploaded... "your photos never leave the phone" is a sales line. > WRONG. Corrected 27 Jul 2026, before this document was ever used to sell anything. This > claim was carried over from RESEARCH.md §8, which was written on 19 Jul and describes an > architecture the app has since moved past. Receipt photos are uploaded.index.html sends > image_base64 to the parse-receipt edge function, which reads the photo with Claude > vision (#168, Mobile M2); on-device Tesseract is now only the fallback path. Receipt images > are also stored, in a private Supabase bucket with a 6-year default retention (#120). > > Marketing "your photos never leave the phone" would have been a false privacy claim on a > compliance product, which is the worst possible place to make one. The real, defensible version > is "your data is held in the UK": the Supabase project is in eu-west-2 (London), verified > 27 Jul 2026. The Anthropic call is a US transfer and needs a documented transfer mechanism, so > it is a compliance obligation rather than a selling point. See > outputs/docs/go-to-market-decisions.md §D1.
Price. We can profitably sit under the anchor (see §4).
Accounts built for how HMRC actually works. Cash basis is HMRC's default for sole traders from 2024/25, which is exactly the model planned. Competitors either ignore accounting or bolt on COGS-heavy inventory tooling built for US Etsy sellers.
Honest weaknesses. Stock control (#33) and food safety logs (#36) are where FoodCore's £55 tier is strong and we are not. We have no brand, no reviews, no customers and no SEO surface (one 520 KB HTML file with no server rendering), while FoodCore is publishing comparison content weekly. And the market has just watched a cheap single-founder tool close, so credibility signals matter disproportionately: a status page, a data export promise, a published backup policy and a real company identity are commercial features here, not hygiene.
4. Suggested Pricing Structure
4a. The model
Trial first, not freemium. Benchmarks put SMB-targeted free-trial conversion at roughly 6 to 10%, against 2 to 5% for freemium, and trials reach monetisation in 12 to 18 days rather than 90 to
Requiring a card raises conversion sharply but suppresses signups, and in a nervous,
price-sensitive market a card wall will cost more than it gains.
Recommendation: a 30-day full-feature trial with no card required, backed by import tooling so a baker can actually get their 70 declarations in during the trial. The trial is only as good as the onboarding, which is why barcode capture is the commercial priority, not a nice-to-have.
When a trial or subscription ends, never lock the data. Drop to a read-only Archive state: everything stays visible and exportable, but label printing and new batches stop. Holding a food business's allergen records hostage is both a bad look and a compliance hazard, and the paid action (printing a label this week) is the one worth gating.
4b. Tiers
All prices in GBP, per business per month, VAT not applicable while under the £90,000 threshold. Annual billing at 10x the monthly price (two months free).
Starter
Standard(recommended)
Pro
Monthly
£9
£19
£39
Annual
£90 (£7.50/mo)
£190 (£15.83/mo)
£390 (£32.50/mo)
Who it is for
Cake shed, market stall, weekend baker
Established home bakery with regular orders
Multi-person kitchen, caterer, farm shop
Ingredients and supplier products
Unlimited
Unlimited
Unlimited
Recipes and products
Up to 25 products
Unlimited
Unlimited
PPDS labels, allergens emphasised
Yes
Yes
Yes
Allergen matrix and export
Yes
Yes
Yes
Recipe and batch costing
Yes
Yes
Yes
Per-batch supplier selection on the label
Yes
Yes
Yes
Second-opinion allergen scan and verification tracking
Yes
Yes
Yes
Barcode capture / OFF import
Yes
Yes
Yes
Offline capture (PWA)
Yes
Yes
Yes
Receipt scanning and purchase capture
-
Yes
Yes
Orders with prices
-
Yes
Yes
Cash-basis accounts and monthly P&L
-
Yes
Yes
Label archive and audit export (EHO pack)
90 days
2 years
Unlimited
Users
1
2
5 (roles: owner, baker, verifier)
Stock control
-
-
Yes
Multiple brands or sites
-
-
Yes
Support
Email, 2 working days
Email, 1 working day
Priority
Deliberate choices worth flagging:
The label engine is in every tier, uncrippled. Never sell a partial allergen label. If someone is paying us anything, the compliance output must be complete, or we own a safety problem.
The paywall runs along "running a business", not "being safe". Orders, accounts, receipts, stock and extra users are the upgrade levers.
Products, not labels, are the Starter limit. Per-label metering punishes exactly the behaviour we want (label everything).
£19 buys at Standard what FoodCore charges £19 for at Essentials, and includes much of their £55 Core tier. That is the positioning: their entry price, most of their top tier, plus the batch provenance nobody else has.
4c. Why these numbers
Cost basis. Fixed costs are roughly £50 to £70 a month (§2c). Per-tenant marginal cost is close to zero: the app is a static file and Supabase Pro's included capacity covers many small tenants.
Payment fee drag (Stripe 1.5% + £0.20, plus Billing 0.7%):
Price
Stripe + Billing fee
Net
Fee as %
£9 monthly
£0.40
£8.60
4.4%
£19 monthly
£0.62
£18.38
3.3%
£39 monthly
£1.06
£37.94
2.7%
£90 annual
£2.18
£87.82
2.4%
The £0.20 fixed component is 2.2% of a £9 charge. Push annual billing hard at the low tiers: it nearly halves fee drag, and it removes eleven chances a year for a card to fail. Offer a founder annual rate rather than a monthly discount.
Break-even. Against a £65/month cost base: about 8 Starter customers, or 4 Standard, or 2 Pro. A realistic mixed book of 20 customers weighted to Standard clears roughly £250 to £300 net a month, which is a real side income and, more importantly, funds the compliance and insurance overhead without subsidy.
Launch tactics.
Grandfather early customers on their signup price permanently and say so. In a market that just lost Bake Diary, "the price you join at is the price you keep" is a trust signal as much as a discount.
Publish the price on the site. FoodCore does; opacity reads as expensive.
No setup fees, no per-label fees, no per-user fees below the tier cap, monthly cancellation. Match FoodCore on all four so none of them becomes an objection.
Revisit pricing only after 20 paying customers. Before that, the data is anecdote.
5. Business Setup and HMRC Guidance
5a. Is sole trader the right structure?
Yes, for now, on the stated facts: single owner, revenue well under the £90,000 VAT threshold, minimal capital, and a preference for the least administrative overhead. Registration is a single online step, there are no filing fees, and accounts stay private.
The one real caveat. A sole trader has unlimited personal liability, and this is compliance software: the failure mode is a customer's mislabelled product, not a lost invoice. That is a genuine rather than theoretical exposure. Mitigate with (a) Terms that state plainly that the food business operator is responsible for their labels, (b) professional indemnity insurance from day one, and (c) a trigger to incorporate. Sensible triggers: the first business customer who asks for a signed DPA with an uncapped indemnity, revenue passing roughly £30,000, or the first employee or contractor.
5b. Setting up "Stella Apps"
Check the name is usable. A sole trader may trade under a business name without registering it, but it must not include "limited", "Ltd", "plc" or similar, must not be misleading, must avoid sensitive words (for example "Royal", "British", "Bank", "Insurance", "Institute", "Chartered"), and must not be confusingly similar to an existing brand. Search the UK IPO trade mark register and Companies House before committing, and check domain and social handles. There is no register of sole trader names, so the protection you get is what you create: consider a UKIPO trade mark in class 9 (software) and class 42 (SaaS) once revenue justifies it. Current IPO fees should be taken from the IPO site at the time of filing [unverified here].
✅ DONE 29 Jul 2026 — registered as a sole trader with HMRC.(Recorded under the home address, deliberately: HMRC holds it on the personal record regardless, publishes no register of sole traders, and its post is deadline-bearing — so routing it through a forwarding service would add risk for no privacy gain. See D-07 and #267.) Original guidance retained below. Register for Self Assessment with HMRC. The deadline is 5 October following the end of the tax year in which trading started (so trading begun in 2026/27 must be registered by 5 October 2027). Register early rather than at the deadline: you need the UTR and Government Gateway account before you can do anything else. Registration is required once self-employment income exceeds, or is expected to exceed, £1,000 in a tax year.
Note the trading allowance. If gross trading income in a tax year is £1,000 or less, there is generally no need to register or report. Above it, you can deduct the £1,000 allowance instead of actual expenses, which is only worth doing if your costs are lower than that.
Pay the ICO data protection fee (Tier 1, £52, or £47 by Direct Debit) before processing live customer data.
Take out professional indemnity insurance (from roughly £8-10 a month) before the first paying customer.
✅ DONE 29 Jul 2026 — Mettle business account opened, which is §6b's primary recommendation. This also settles accounting (§6c): FreeAgent now comes free with the account, so the £0 accounting line in §3 holds and no separate decision is needed. Remaining habit worth adopting: a separate savings pot taking 25-30% of every payout on the day it lands. (Original: open a separate business bank account — see §6. Not legally required for a sole trader, but it is the difference between an hour and a weekend at tax time.)
Publish the legal stack before launch: Terms, Privacy Policy, DPA, sub-processor list, and contact and complaints details.
Keep Stella Apps separate from the bakery. If Whisk & Whimsy is a tenant of the product, document it as such. Mixing the two sets of books, or the two sets of customer data, creates both a tax and a data-protection problem that is trivial to avoid now and awkward to unwind later.
5c. Tax and compliance checklist
Item
Detail
Timing
Self Assessment registration
By 5 October after the end of the first trading tax year
Do it at launch
Unique Taxpayer Reference (UTR)
Issued after registration
Allow weeks, not days
Tax return and payment
File and pay by 31 January after the tax year ends; late filing penalties start at £100
Annual
Accounting basis
Cash basis is the default for sole traders from 2024/25: income when received, expenses when paid. (eligible up to £150,000 turnover) ⚠️ corrected 5 Aug 2026 — the £150,000 entry / £300,000 exit thresholds were abolished from 6 Apr 2024; cash basis is now available to an unincorporated business of any size, and accruals is the opt-out. Same correction already carried in RESEARCH.md §7. See §6c for where the basis is actually set.
Ongoing
Class 4 NIC
6% on profits between £12,570 and £50,270; 2% above
Via Self Assessment
Class 2 NIC
Abolished from April 2024. Profits above the Small Profits Threshold (£7,105 in 2026/27) get State Pension credit automatically; below it, voluntary Class 2 is £3.65/week if you want to protect your record
Consider annually
Record keeping
Keep business records to support the return, per HMRC guidance (self-employed records are generally kept for 5 years after the 31 January filing deadline)
Ongoing
VAT
Register once rolling 12-month taxable turnover exceeds £90,000 (frozen for 2026-27), or if you expect to exceed it in the next 30 days
Check monthly
VAT on overseas B2C digital sales
Place of supply is the customer's location, with no threshold overseas
Only if you sell outside the UK
MTD for Income Tax
Live from April 2026 for qualifying income over £50,000; over £30,000 from April 2027; over £20,000 from April 2028, assessed on 2026-27 qualifying income
The £20,000 step is low enough to catch a successful side business. Keep digital records from day one so it is a non-event
ICO fee
£52/year (£47 DD), renewed annually
At launch
Insurance
Professional indemnity, plus consider cyber
At launch
UK GDPR
Privacy notice, record of processing, DPA, sub-processor list, DSAR and breach process, complaints route (DUAA duties from 19 June 2026)
Before first customer
6. Banking and Accounting Recommendations
6a. Tide, assessed against this specific business
What it costs and gives you (verified July 2026): free plan with no monthly fee and no minimum balance, a Mastercard debit card, and invoicing built into the app. Reported FSCS protection to £120,000 on new accounts [verify at application; FSCS protection normally depends on the underlying deposit-taking bank]. Accounts commonly open within hours. Direct Sage bank feed, plus Xero, QuickBooks and FreeAgent integrations on the free plan.
Charges: 5 free transfers a month then 20p per transfer in or out; £1 per ATM withdrawal; 3% on cash deposits at PayPoint; 90p per cheque deposited by app; 2.75% FX on overseas card use on the free plan. No overdraft.
Verdict for Stella Apps: perfectly workable, and the transfer fee barely bites. A SaaS business has a handful of movements a month (a Stripe payout, a few supplier Direct Debits, a transfer to a tax pot), so 20p per transfer is pennies. The 3% cash and 90p cheque fees are irrelevant to a business that will never handle either. The weak points are the absence of an overdraft and the FX rate, and FX only matters if you buy in USD (Supabase and several of the tools above bill in dollars, so a card with better FX is worth something).
6b. Alternatives worth taking seriously
Provider
Monthly fee
Notable for
Watch out for
Mettle (NatWest)
Free
Includes the full FreeAgent accounting platform free, with every feature, for as long as you hold the account (requires at least one transaction a month). FreeAgent is HMRC-recognised for MTD
No cash deposits (irrelevant here); app-first
Starling
Free
A full UK bank rather than an e-money firm; no per-transfer fee; cash deposits at the Post Office; overdrafts available; Xero, QuickBooks and FreeAgent integrations
Sole trader accounts are subject to eligibility
Monzo Business Lite
Free
Unlimited free UK transfers; Pro tier at £9/month adds invoicing and virtual cards
Fewer integrations than Starling or Tide
Revolut Business / ANNA
Free tiers plus paid
Multi-currency, useful if USD tool spend grows; ANNA layers on tax help
No cash deposits; paid tiers add up
Tide
Free
Best-in-class invoicing inside the banking app; very fast onboarding
20p transfers after 5/month; £1 ATM; no overdraft
Recommendation.
✅ Primary: Mettle — CHOSEN AND OPENED 29 Jul 2026. Precisely because it makes the accounting software free. FreeAgent is a £150-300 a year product being given away with a free bank account, it is HMRC-recognised for MTD, and it removes an entire line item and an entire decision.
Secondary: Starling or Tide. Starling if you value being with an actual bank with an overdraft option; Tide if the built-in invoicing genuinely suits how you bill. Holding a second account at a different institution is cheap insurance against a frozen account, which is a real and well-documented risk with app-based providers. ⏸ Deferred 5 Aug 2026 (#270): revisit when Stripe goes live (#206), not now. Little point insuring an account with nothing in it, and a dormant business account risks being closed anyway. ⚠️ The accepted trade-off: opening one takes days-to-weeks, so this may end up done under time pressure — start the application in parallel with #206 rather than after it. Mettle stays primary regardless, since FreeAgent is free only while it sees at least one transaction a month.
Set up a separate savings pot and move 25-30% of every payout into it for tax on the day it lands. This is the single highest-value habit for a new sole trader.
6c. Accounting: simple and cheap
Option
Cost
Fit
FreeAgent via Mettle / NatWest / RBS
£0
Recommended. Full platform, no feature limits, MTD-ready, bank feed included
QuickBooks Sole Trader
£10/month
The cheapest mainstream paid option with solid MTD coverage
Crunch (Free / Pro / Premium)
£0 / £10+VAT / £21.60+VAT per month
Premium bundles a named accountant, worth considering in year two
Xero Ignite
£16/month
More than this business needs
Pandle, QuickFile, Zoho Books, Wave
Free tiers
Viable if you stay off Mettle, but check MTD readiness before relying on one
Accountant for the annual return
Typically a few hundred pounds a year [unverified: varies widely]
Worth it in year one to get the treatment of software costs, home-office use and the Stella Apps / bakery separation right. ✅ Decided 5 Aug 2026 (#270): engage before the 5 Apr 2027 year end, not at filing time in Jan 2028 — those questions change how you record during the year, and an accountant engaged after year end can only tidy, not steer. Add cash-basis loss relief to the list: no longer carry-forward-only since 2024/25
Practical bookkeeping notes for this business:
Book gross revenue and record Stripe fees as an expense. Booking net understates turnover, which matters for the £90,000 VAT test.
Categorise recurring tool spend (Supabase, Netlify, email, insurance, ICO) as a single "software and subscriptions" category so the monthly cost base is visible at a glance. ⚠️ Superseded 5 Aug 2026 (#270), twice — and the second attempt was also wrong, so both are recorded here. The original list was stale (Netlify has been gone since 3 Aug) and one lumped line destroys the fixed-vs-usage split #216 exists to measure. The second attempt proposed creating three custom categories, possibly split across FreeAgent's Cost of Sales and Admin Expenses tabs. Looking at the actual category list killed that too: FreeAgent already ships everything needed — create nothing. Two things the real UI settles:
The dropdown is labelled "Tax Return Box", offers thirteen fixed options, and cannot be added to. A category selects a box; it never defines one.
None of the thirteen is "cost of goods bought for resale" — that box belongs to the Cost of Sales tab, whose categories are all locked to a single "Cost of Goods" box. For a software business with no physical goods, use Admin Expenses only and skip Cost of Sales. ⚠️ Amended once the Cost of Sales tab was actually read: it holds 101 Cost of Sales, 102 Commission Paid, 103 Materials, 104 Equipment Hire, 150 Subcontractor Costs. Three are goods-oriented and useless to us, but 101 and 102 are exactly the two lines a SaaS business needs, so Cost of Sales is in use after all — for those two rows only.
The mapping, existing categories only (also kept as a one-page reference at outputs/docs/freeagent-category-mapping.pdf):
Cost of Sales — the per-sale costs: 🎯 Anthropic API → 101 Cost of Sales; Stripe fees → 102 Commission Paid.
Admin Expenses — fixed overhead: Supabase (both projects), Cloudflare and the domains → 268 Web Hosting; Google Workspace and Resend → 269 Computer Software; insurance → 364; ICO fee → 361 Subscriptions; trading address (#267) → 250 Office Costs (a service, not property, so not 251 Rent); accountant → 292; home working → 366 Use Of Home.
🎯 The point of putting those two in Cost of Sales is that FreeAgent then computes Income less Cost of Sales = gross profit by itself — which is exactly #216's output. Anthropic is the only genuinely per-customer cost in the estate, so it is the numerator; leave it merged into hosting and that number has to be reconstructed by hand from records that no longer distinguish it.
Supabase deliberately stays on 268, not Cost of Sales. Textbook SaaS accounting puts hosting in COGS, but ours is a flat per-project fee regardless of customer count — fixed overhead, not cost of sale. Move it to 101 if Supabase billing ever becomes usage-based.
⚠️ For a service business with no goods, "Cost of Goods" is a slightly odd SA103 box. Immaterial below £90k (single total figure) — on the list for the year-one accountant, not a blocker.
USD charges should be recorded at the rate actually paid, not a notional rate. Keep the card statement.
Under cash basis, most equipment is simply expensed in the month it is paid for, so a flat "expense it when it leaves the account" model is correct.
Where the accounting basis is actually "set" (added 5 Aug 2026, #270). There are three places it could plausibly live and only one of them is an action:
HMRC — nothing to do. There is no form, notification or online setting for the basis. Since 2024/25 cash basis is the default, so it applies by operation of law from the first day of trading unless you actively opt out. Registering as a sole trader did not ask, and should not have.
The Self Assessment return — a non-action to protect. Cash basis = leave SA103S box 8 ("If you used traditional accounting rather than cash basis, put 'X' in the box") blank. The only way to get this wrong is to tick it. First relevant filing is the 2026-27 return, due 31 Jan 2028.
FreeAgent — the one real action. FreeAgent asks for the accounting basis alongside the accounting dates during account setup, and it drives every report and its Self Assessment area. If it was left on accruals, the books disagree with the return. Switching needs level-8 (full) access, is unavailable to a user holding the 'Accountant' role, and becomes unavailable once a return has been filed through FreeAgent on the other basis — which is why this is cheap now and expensive later.
Cash basis losses are no longer trapped: the pre-2024/25 restriction to carry-forward-only was removed at the same time as the turnover thresholds, so a first-year loss can go sideways against general income or be carried back. Worth raising with the year-one accountant rather than self-serving, since it interacts with the Stella Apps / bakery separation.
⚠️ Do not confuse the accounting basis with the VAT cash accounting scheme — different regime, different election, and irrelevant here while D-16 holds (deliberately not VAT-registered).
Keep digital records from day one. Even if MTD for Income Tax does not apply yet, the £20,000 threshold arrives in April 2028 and is assessed on 2026-27 income.
Estimated year-one running cost for Stella Apps:
Item
Annual
Supabase Pro ($25/month) — ⚠️ two projects since #137 (production + staging); take the real figure from the invoice, not from this table
~£240, plus staging
NetlifyCloudflare Workers
~~£0 (free tier) to ~£180 (Pro)~~ £0 — static-asset requests free and unlimited, deploys unmetered (D-20)
Transactional email
£0 (Resend free tier)
Anthropic API (receipt + recipe parsing) — ⚠️ missing from this table until 5 Aug 2026; it is the only genuinely usage-scaling line, so it is the one that matters most to #216
usage-based [unverified — read the actual bill]
Domains (provenbatch.co.uk, stellaapps.com)
~£15 each
Google Workspace (support mailbox, D-21)
[unverified — read the actual bill]
ICO data protection fee
£52 (£47 by DD)
Professional indemnity insurance
~£100-150
Banking and accounting
£0 (Mettle + FreeAgent)
Total
~~~£410 to £640~~ ⚠️ stale as of 5 Aug 2026 — do not quote it. Netlify has gone to £0, but a second Supabase project, the Anthropic API and Google Workspace all arrived after this was written, and two of those are unverified. Rebuild this from FreeAgent's own figures once the categories above have a full month of data — which is the whole point of doing the categorisation now rather than reconstructing it later
Against the tiers in §4, that is covered by roughly four Standard customers.That customer count rests on the total above and is stale with it — recompute both together, once, from real invoices. Recorded as the first concrete input to #216.
Open questions this report does not close
Would a baker actually pay? Still unvalidated. Two home bakers who are not us remain the cheapest falsification of everything above (carried over from RESEARCH.md §7).
Precise UK market size. The FSA does not publish a home-based food business total. Still open.
Planglow's pricing could not be verified.
Whether "Stella Apps" is free to use as a name and mark. Needs an IPO and Companies House search before any spend on branding.
The single-tenancy blocker (#48) gates all of it.Wrong, corrected 27 Jul 2026: #48 shipped in v0.14.0. The real gates are billing, a separate staging environment (#137), the legal stack and support tooling. See outputs/docs/launch-plan.md for how they are sequenced.
Full report: outputs/docs/gtm-channels-research.md (ten sections: TAM/SAM/SOM, online communities, social & creator channels, SEO & content, paid acquisition, email + the PECR/DUAA legal rails, events calendar, partnerships, a phased channel plan mapped to the launch timeline, and follow-ups). Six parallel research passes; almost all evidence is search-extract (the proxy 403'd nearly every direct fetch — caveat stated once at the top of the report). Headlines worth having here:
Market size, finally bounded (answers old open question 1). UK register ≈ 540k rated / 605k total food businesses (the 430k figure elsewhere was the older rated-only count — corrected in place). TAM = 605k × an assumed 25–40% doing PPDS ⇒ 150–240k businesses, £27–43m ARR at our blended £15/mo. Reachable SAM ≈ 110–190k (home-based businesses modelled at 60–120k active — the FSA publishes no stock total). Realistic 3-year SOM: 1,000–5,000 subscribers = £180k–£900k ARR. Honest read: a lifestyle-to-small-business outcome, not a venture market, and Bake Diary dying at €6.95/mo means the £9 tier alone can't carry the business — the 30%+ on £19+ mix assumption is load-bearing for viability.
🔴 Cold email to most of our audience is illegal. Sole traders are individual subscribers under PECR — consent or soft opt-in only; the DUAA (in force 5 Feb 2026) left that unchanged and raised maximum PECR fines from £500k to £17.5m/4%. ICO cases are lost on missing consent records → consent log (timestamp, source, wording) in Supabase from day one. DUAA's one gift: first-party analytics cookies are now banner-free (lost if data feeds advertising) — so Plausible-class analytics, no ad pixels, which also suits the brand: a compliance product must not open with a tracking banner.
Paid acquisition is not a primary channel at £9–39/mo. SMB SaaS CAC norms are 2–7× our £30–100 allowable CAC. The one paid bet: exact-match Google Search on "allergen label software"/"natasha's law software" at £5–10/day — SERPs show SEO investment but no evident ads (verify in Ads Transparency Center — dashboard-only from a cloud session).
Every high-value UK baker Facebook group is a coach's funnel (Annie Bennett's Home Baking Business Community UK ~7k, A Bigger Slice, The Cake Business Club, CakeFlix's gated groups). The route in is an admin partnership — guest allergen-labelling masterclass, affiliate cut — not posting. Confirms and extends D-17. Non-baker verticals route through associations: NCASS (⚠️ channel-conflict risk — it bundles its own compliance offering), NMTF (clean, lead here), FRA, craft-butcher bodies. Compliance questions get asked on Mumsnet, r/AskUK, and in the groups.
SEO flank is open where it matters: FoodCore double-dips the commercial SERPs but ignores the home-baker vertical, farm shops (that SERP is literally Wikipedia), and batch-level traceability — our differentiator story, unclaimed except one blog. "Bake Diary alternative" is the fastest win: orphaned UK/Irish users, every current ranker US/inventory-first, none mention PPDS. Directory listings (Capterra/GetApp/G2/Serchen/SourceForge) are free, rank by themselves, and carry a ~3× AI-citation multiplier; Reddit is the #1 source AI search engines cite. Expect near-zero Google traffic before month 4 of the site. ⚠️ Correction of consequence, 9 Aug 2026: the Reddit finding still stands as research, but Dave dropped Reddit as a channel — so the AI-citation upside it describes is one we have decided not to collect. The directory listings, which carry the ~3× multiplier and cost nothing, are now the whole of this lane. ✅ REVERSED 10 Aug 2026 — the correction above lasted one day. Dave reinstated Reddit and the account exists (u/ProvenBatch, secured; see gtm/social-accounts.md). The AI-citation upside is back on the table, so this lane is directory listings and Reddit again, not directories alone. UK Business Forums stays dropped. Kept rather than rewritten because the sequence is the point: a finding survived two reversals of the decision built on it, which is the argument for recording research separately from the choices made about it. 🔻 FINAL, 15 Aug 2026 — the chain ends here: Reddit is dropped PERMANENTLY (Dave; the account is being deleted, Inoreader with it — both generated work he doesn't have time for). The Reddit finding above still stands as research; the upside is deliberately not collected, and directory listings are the whole of this lane — the 9 Aug correction turned out to be the durable one. Related finding, verified same day against Reddit's own help centre: Reddit closed self-serve API app creation in Nov 2025 ("Responsible Builder Policy" — manual approval, personal-use requests effectively never granted), so automation on the channel was already impossible without an approval Reddit rarely gives. Do not re-add the channel; this supersedes both reversals.
Creators are nano-scale and already commercial: UK cake-business coaches at 1–10k followers, £20–150/post, with affiliate precedent (CakeFlix 30%, Craftybase 12-month rev-share) → a recurring-commission ambassador scheme fits our price point. YouTube's "Natasha's Law" inventory is stale 2021 institutional video — a current maker-voiced series has no incumbent. The Jun 2025 TikTok Shop allergen scandal is a standing content hook.
Events: visit, never exhibit (£3–5k/stand vs ~£300 LTV; ROI literature says self-serve low-price products don't pencil). Cake International, NEC, 6–8 Nov 2026 — beta month, an hour away — walk it with beta invite cards. Pre-beta: Speciality & Fine Food Fair (8–9 Sep, ⚠️ date conflicts in listings) and Street Food Business Expo (29–30 Sep, NCASS runs the keynote theatre).
Partnership rails already exist: High Speed Training runs a public affiliate programme; councils host third-party home-baker guidance (West Norfolk's cake-makers PDF is the precedent for an EHO-signposting ask, start with Buckinghamshire); insurers' cover assumes compliance (CMTIA reachable founder-to-founder); a wholesaler (Pilgrim Foodservice) already markets a Natasha's Law solution to its customers.
Full write-up
From outputs/docs/gtm-channels-research.md
ProvenBatch — Go-to-Market Channels & Market Sizing Research
Researched 31 Jul 2026, six parallel research passes (market sizing · online communities · social & creator channels · SEO & content · paid + email/automation · events & partnerships). Companion to stella-apps-business-research.md (what the market looks like), go-to-market-decisions.md (what we have decided) and launch-plan.md (when things happen). This document says where the customers are and how to reach them. It builds on decisions already made — D-24 (small food businesses generally, never bakery-first), D-17/#218 (beta via Facebook groups + local intros), D-23 (Astro site for SEO), D-02 (£9/£19/£39), D-01 (UK only) — and nothing here contradicts them.
⚠️ Evidence caveat, stated once for the whole document. The agent proxy 403'd essentially every direct page fetch during this research (food.gov.uk, legislation.gov.uk, facebook.com, foodcore.io, ICO guidance pages, YouTube watch pages, ukbusinessforums.co.uk, farmretail.co.uk and dozens more). Nearly every claim below is search-extract evidence — consistent across sources where possible, flagged individually where thin. Good enough to plan against; anything that gates a spend or a legal position should be confirmed by eye first. Individual "[extract-only]" flags are kept where the claim is load-bearing or contested.
0. Headlines — what this research changes
Market size is finally bounded (§1). The UK register is ~540k rated / ~605k total food businesses (larger than the 430k in earlier docs — that was the older rated-only count). TAM for PPDS-obligated businesses ≈ 150k–240k businesses (£27m–£43m ARR); reachable SAM ≈ 110k–190k; a realistic 3-year SOM is 1,000–5,000 subscribers = £180k–£900k ARR. Honest framing: a lifestyle-to-small-business outcome, not a venture-scale market — and two load-bearing numbers (the PPDS fraction, the active home-business stock) are still modelled, not known.
Cold email to most of our audience is illegal (§6). Sole traders are individual subscribers under PECR, so "B2B cold email is fine" is false for home bakers and market traders. The DUAA (in force 5 Feb 2026) raised PECR fines to UK-GDPR level (£17.5m/4%). Consent-logged, opt-in email only.
Paid ads are not a primary channel at £9–39/mo (§5). SMB SaaS CAC norms (£150–550) sit 2–7× above our allowable CAC. The one paid bet with a real chance: exact-match Google Search on "allergen label software / natasha's law software" at £5–10/day.
Every high-value baker Facebook group is a coach's funnel (§2). The route in is a partnership with the admin (guest masterclass, affiliate cut), not posting. D-17's "participate first" rule is confirmed and extended.
The SEO flank is more open than FoodCore's presence suggests (§4). FoodCore double-dips the commercial SERPs, but the home-baker vertical, farm shops (that SERP is literally Wikipedia), batch-level traceability (our differentiator), and "Bake Diary alternative" (orphaned UK/Irish users, no ranker mentions PPDS) are all winnable now.
The creator market is nano-scale and already monetised (§3). UK cake-business coaches run 1k–10k followings and already do affiliate deals (CakeFlix 30%, Craftybase 12-month rev-share as templates). YouTube "Natasha's Law" content is stale 2021 institutional video — a current, maker-voiced series has no incumbent.
Visit events, don't exhibit (§7). Exhibiting doesn't pencil at our price point (£3–5k/show vs ~£300 LTV). Cake International, NEC, 6–8 Nov 2026 — the densest gathering of home cake makers in the UK, in our beta month, an hour from Aylesbury — is the one to walk with beta invites.
Partnership rails already exist (§8). High Speed Training runs a public affiliate programme; councils demonstrably host third-party guidance for home bakers; insurers sell policies whose cover assumes compliance; a wholesaler (Pilgrim Foodservice) already markets a "Natasha's Law solution" to its customers. These are conversations, not builds.
1. Market sizing — TAM / SAM / SOM
1a. The register (2026 baseline)
Figure
Value
Source
FHRS establishments with a numeric rating (England/Wales/NI)
⚠️ Correction to earlier docs: RESEARCH.md §2/§11 and the decision register's §E carry "430,000+ FHRS businesses" — that appears to be an older rated-only count. Use ~540k/~605k from here on (corrected in RESEARCH.md §12).
Registration flow: 90,613 new registrations 2023/24 (vs 83,594 in 2022/23), with 20% growth in home-caterer and restaurant/takeaway sectors; H1 2024/25 ran at the same ~90k/yr pace; >39,500 new registrations were still awaiting first inspection as of Feb 2025 (FSA board papers via gov.uk mirror + Food Safety News, extract-only). ~85–90k gross new registrations a year against a slow-growing register implies very high closure churn among new food businesses — which feeds the SOM churn assumptions below.
1b. Segment counts (IBISWorld enterprise counts unless noted, all extract-only)
Bread & bakery production 3,002 (2025) · bakery cafes 1,793 · bakery product retailing 4,310 (2026) · butchers 5,588 (avg 5.5 employees) · delicatessens 1,509 (2026) · cafes & coffee shops 8,403 (plus ~12,400 independent coffee outlets, Lumina 2024) · farm retailers ~1,580 (Harper Adams 2022; FRA cites 1,581) · food & drink manufacturers 12,195, 99% SME (FDF 2025) · new takeaway companies 2023–25 67,293 (Companies House via h2products, soft). No chocolatier or preserve-maker counts were obtained (gap).
1c. Home-based food businesses — modelled, not published
37% of 92,540 new registrations since Mar 2020 were from domestic kitchens (FSA campaign via Food Safety News, Mar 2022) — confirms and dates the figure already in RESEARCH.md §2.
The FSA still publishes no stock total of home-based businesses — re-confirmed this pass.
Modelled estimate: ~440–460k gross registrations Mar 2020→mid-2026 × ~35–37% domestic ≈ 155–170k gross home-based registrations; applying 40–60% cessation within two years (assumption, justified by the register's slow net growth) ⇒ active stock ≈ 60,000–120,000, central ~90,000. A modelled range, not a fact — the two follow-ups that would replace it with a fact are in §9.
1d. What fraction does PPDS — the load-bearing unknown
The 2019 PPDS Impact Assessment (legislation.gov.uk/ukia/2019/144) would state the official in-scope count; it 403'd on every route and extracts didn't surface the number. Honest gap. What the July 2023 FSA evaluation of Natasha's Law does say (via FSA blog / British Baker extracts): 91% of PPDS businesses aware of the law, 68% say they have the information they need, 26% changed what they sell — 17% moved food out of PPDS entirely, 16% shifted to buying in prepacked; trading standards testing found large minorities in breach (one headline: "half of businesses tested"; FSA 2023–24 testing 17 of 47 items with unlabelled allergens — weakest provenance, verify before quoting). Working assumption: 25–40% of registered food businesses regularly produce PPDS. Everything downstream is sensitivity-tested, not known.
SAM ≈ TAM deliberately: our target list is the whole PPDS-doing micro population — the excluded segments (chains, institutions, mid-size manufacturers) were never counted in.
SOM constraints: (1) only ~48% of micro businesses adopt technology at all (gov.uk/OECD, definition unverified); (2) category precedent — Craftybase "10,000+ makers" lifetime, global; Bakesy "74K+ users" global, free-inflated; Bake Diary "thousands" of UK/Irish bakers at €6.95/mo and it still closed (May 2025); (3) churn arithmetic — at 4%/mo (mid-range SMB benchmark, 3–7%/mo across venasolutions/kalungi/churnkey), steady state = 25 × monthly gross adds, so 40–120 organic adds/month ⇒ 1,000–3,000 subscribers. Anything claiming ≥10% of SAM should be rejected as overclaiming.
Sensitivity: the PPDS fraction halves TAM if halved; the home-based stock dominates SAM; churn at 7%/mo cuts the subscriber base to ~57% of the 4%/mo case. Bake Diary's death at €6.95/mo suggests the £9 tier alone cannot carry the business — the 30%+ on £19+ mix assumption is load-bearing for viability, not just sizing.
2. Online communities — where the owners actually are
2a. The structural finding
Every high-value UK home-baking group is somebody's funnel. The "community" is 4–6 free Facebook groups, each run by a coach selling courses to exactly our ICP:
Community
~Size
Structure
The Home Baking Business Community UK (Annie Bennett)
7,000+ [extract-only]
Feeder for her paid "HBP Society"; guest experts deliver masterclasses inside the paid tier
A Bigger Slice – The Cake Collective
not visible
Free group atop paid cake-business classes; "highly engaged" per third parties
The Cake Business Club (Kate Tynan)
small, pure ICP
Free community + paid coaching
Business of Cake
new
Learning community; also sells a cake pricing calculator
CakeFlix members' groups
subscriber-gated
210k site members claimed [extract-only]; access only via commercial partnership
Implication: don't fight the funnel — join it. Offer admins a free allergen-labelling masterclass (their members' #1 anxiety, our exact ground), an affiliate cut, or free accounts for their paid tiers. One admin partnership beats a thousand promo posts, and it is the only reliable way past their gates. This is how D-17's "participate first" plays out in practice.
2b. Other verticals route through associations, not groups
Street food / caterers: NCASS (5,000+ members, 11,000+ outlets, ~90% street food/mobile/events [extract-only]) — no public forum; the channel is supplier membership (newsletter inclusion, "approved supplier" badge, discounted ads). Trader groups exist on Facebook (Street food/Food Trucks/Traders/Equipment U.K; Street Food Traders – Mobile Caterers) but peer intel moves through WhatsApp groups run by market organisers — the way in is via the organiser, not the group.
Market traders: NMTF (~25,000 traders represented, £150/yr sole-trader membership) + a cluster of FB groups (Events Organisers and Market Traders; Markets & Events UK; Connecting Stall Holders & Event Organisers UK).
Farm shops: no large peer FB group exists; the vertical is the Farm Retail Association (~325 members; the wider population is 1,000+ farm shops/markets).
Butchers: Craft Butcher International (1,400+ members as of Dec 2020, likely larger), Butchers Collective (supplier-run), National Craft Butchers / Scottish Craft Butchers.
Chocolatiers/preserves/cafes: UK Chocolatiers FB group; Guild of Jam and Preserve Makers; Coffee Shop Owners UK + Coffee Forums UK's Business Owner subforum.
No owner-side food-hygiene discussion group exists — compliance talk happens inside the trade groups above, which is exactly where our answers belong.
2c. Forums, Reddit, and where the compliance questions get asked
🔻 FINAL, 15 Aug 2026 (Dave): Reddit is DROPPED PERMANENTLY — do not re-add. The 10 Aug reinstatement is reversed for good; u/ProvenBatch is being deleted. Dave's reasoning: the channel generates work he doesn't have time for. A contributing fact found the same day: Reddit closed self-serve API app creation in Nov 2025, so tooling could not have carried the load either. The §2c answer lane and the §4e AI-citation upside are deliberately not collected via Reddit — Mumsnet, the Facebook groups and the directory listings carry those lanes instead. The research below is kept per this file's rule: it is the priced record of what the decision gives up, not a recommendation. The 9 Aug decision it re-confirms:
🔻 DECISION, 9 Aug 2026 (Dave): Reddit and UK Business Forums are dropped as channels — neither account is being opened. The research below is deliberately kept, not deleted. It is not wrong; it is the evidence for what the decision costs, and without it the next reader re-derives "we should be on Reddit" from scratch. ⚠️ The cost is concentrated in one line: Reddit is the #1 source AI search engines cite (§4e), so the AI-citation compounding is what is being given up. Mumsnet and the Facebook groups are unaffected and carry the answer-content lane instead.
UK Business Forums (250,000+ members claimed) — the best-codified promo rules found anywhere: self-promo only in the Marketplace section, only for paid Business Members; unprompted DMs explicitly banned. A legitimate paid promo lane + evergreen SEO-indexed answer threads.
Reddit is a Q&A/SEO channel, not a promo channel: r/AskUK (2.4M) is where "can I sell food from my kitchen?" gets asked as life-admin; r/smallbusinessuk exists; all strictly no-promo. Notably, Reddit is the #1 source AI search engines cite (§4e) — honest answers there compound.
Mumsnet Talk is the richest open Q&A venue found: "How do you start selling cakes from home?", "Is it possible to make a living as a home baker?", threads already citing the 28-day registration rule and allergen training. Exactly ProvenBatch's answer territory.
The "do I need allergen labels for my cakes" SERP is currently won by council PDFs (East Suffolk, Reading, West Norfolk, Torbay) and one competitor-adjacent site (Allergen Checker) — a direct content gap (§4).
Write off: Discord (nothing there), LinkedIn groups (wrong audience — B2B/supplier flavoured), giant global cake-decorating groups (hobbyist reach, no targeting), The Staff Canteen (employed chefs, not owner-operators).
2d. Ranked shortlist (10)
The Home Baking Business Community UK — admin partnership (Annie Bennett)
A Bigger Slice – The Cake Collective — admin partnership
UK Business Forums — paid Marketplace lane + answer threads 🔻 dropped 9 Aug 2026 (and still dropped after Reddit's 10 Aug reinstatement — the paid-membership reason is unaffected)
Open item: actual member counts for every FB group need a logged-in Facebook session (~30 min of Dave-side verification).
3. Social & creator channels
3a. The creator landscape is nano-scale and already commercial
UK cake-business coaches (Kate Tynan/The Cake Business Club ~7.4k IG, Karen MacFadyen ~1.2k, Lauren Smy ~0.9k, Emma Stewart, @cakeandbusiness — all [extract-only]) run small but 100%-ICP audiences, and the tier already does software collabs (Castiron, a US baker-software firm, has used Tynan). CakeFlix (52k YouTube subs, 210k members claimed) runs an affiliate programme at up to 30% commission; Craftybase pays a recurring commission for the referred customer's first 12 months. Finch Bakery (157.7k TikTok) — not a coach, but the UK's loudest "bakery small business" voice — has already posted an allergen-labels video citing Natasha's Law: direct proof the topic plays on baker TikTok.
Rates: UK nano (1–10k) £20–150/post; micro (10–100k) £80–1,500/post; video +25–50% (Whito UK 2026 benchmark [extract-only]). At our price point, flat fees only make sense at nano rates — the fit is a recurring-commission ambassador scheme for coaches and podcast hosts, topped with £50-tier podcast sponsorships.
3b. Platform findings
TikTok/Instagram: the audience self-identifies via #cakebusiness, #uksmallbusiness, #ukcakemaker and day-in-the-life discover pages per trade; TikTok claims 1.5M UK businesses on platform; Swansea Council's market-trader TikTok drew 1.1M views in a year. "Bakesy app review" is an established TikTok search genre — home bakers search app names + "review/how to use", which is how Bakesy (US) grew: organic, non-sponsored recommendation videos from real bakers. The June 2025 TikTok Shop allergen scandal (BBC investigation: sellers listing "spices" as allergen info) is a standing stitch/duet hook — and nobody owns the food-safety-educator seat on UK TikTok.
YouTube: the "Natasha's Law explained" inventory is almost entirely 2021–22 launch-window content by institutions and vendors (FSA, councils, Brother), not creators bakers follow. View counts were unobtainable (watch pages 403), so the gap is real but unquantified — validate with the first two videos before scaling.
Street food / farm shop: no creator equivalent exists; the SERP is owned by SaaS/insurer content marketing (Square, Toast, Simply Business, High Speed Training) and the institutional voices are NCASS and Speciality Food Magazine (ABC-certified 8,498 circulation, farm shop/deli readership).
Podcasts: The Business of Cake Making Podcast (156+ episodes, UK cake makers); My Baking Journey (OLBAA) sells sponsorship "from as little as £50" [extract-only] — the cheapest possible test of podcast spend; NCASS newsletter/podcast via supplier membership; Froghop Food Founders / Food Talk Show for artisan producers. Skip broad UK-entrepreneur podcasts (wrong precision per pound).
3c. Where a solo founder's effort goes first
TikTok + IG Reels, one asset set (daily-ish, £0) — the 75-second label-derivation demo is the native genre; post the same vertical video to both.
YouTube search-intent long-tail (1–2 videos/month, compounds) — own "allergen labels for cakes UK" the way FoodCore is trying to own the written SERP.
Niche trade media & bodies (a few emails, £50–low hundreds) — covers the non-baker verticals TikTok reaches less well.
4. SEO & content
4a. Keyword universe (SERP-density evidence; no volume tool was reachable — do a free Keyword Planner pull before writing)
(a) Compliance questions — "natasha's law" (FSA unwinnable at #1; commercial-adjacent variants winnable), "PPDS labelling" (council pages are thin PDFs — beatable below the FSA), "14 allergens list" (people pay £2–4 on Etsy for posters the FSA gives away — demand + weak supply), "do I need allergen labels" home-baker phrasing (rankers are cake-supply shops and hobby blogs, no software vendor, FoodCore absent).
(b) Task keywords — "allergen matrix template" (crowded but all static downloads; rival Paddl already runs the template-magnet playbook; nobody offers a generated, always-current matrix), "recipe cost calculator" (very high demand, US-heavy, nobody ties costing to labelling), "ingredients list generator" (rankers are US/FDA-first with a UK skin — a UK-native generator with bolded allergens is a gap).
(c) Commercial — "food labelling software UK" / "allergen label software UK": FoodCore holds both a money page AND the comparison listicle (classic double-dip, self-ranking #1 in its own comparisons); other rankers (Nutritics, Kafoodle, IndiCater, InstaLabel) sell to hospitality chains, not micro-producers — positioning gap at the £9–39 end. "Bakery software UK" is directory-dominated (win by being listed, not by blogging). "PPDS label printer": every ranker sells hardware; the software-first "your £30 thermal or inkjet is fine" answer is absent.
(d) Vertical long-tail — the open flank — cake makers (no vendor owns it), street food (NCASS member-walled; FoodCore has a thin LP), farm shop (the SERP is literally Wikipedia — zero vendor content), butchers, preserves (SERP is all Canva/Avery templates), wedding cakes (only a law firm and a YouTube video rank). Batch-level traceability — our differentiator story — is unclaimed except for one blog (CraftBatch); FoodCore's allergen tracking is recipe-level.
4b. Competitor content operations
FoodCore runs a real two-layer operation: programmatic vertical landing pages (/food-labelling-software, /natashas-law-labelling-software, /market-stall-food-labelling-software, /recipe-costing-software-for-bakeries…) plus self-serving blog listicles ("Allergen Label Software UK", "Best Kitchen Management Software UK 2026"), and is listed on Capterra/GetApp/G2/Software Advice/SaaSworthy. What they don't cover (observed absences): home-baker compliance questions, wedding cakes, farm shops, butchers, preserves, batch-level traceability, enforcement stories, free interactive tools. Their content is comparison-shaped, not fear/answer-shaped. Planglow targets cafés/food-to-go (labels+software bundle; Erudus data partnership); Nutritics is enterprise/hospitality; smaller players to watch: Paddl (template magnets), LabelFood (city programmatic pages), CraftBatch (batch tracking), InstaLabel, Allergen Checker — and, added 24 Aug 2026, AllergenKit (twelve SEO guides on exactly our buyer's keyword set — natashas-law, selling-food-from-home, 5-star-hygiene-rating — plus live free tools; Dave found them organically, so the funnel demonstrably works; full analysis the master competitor sheet).
Content gaps no vendor answers well: "is this specific thing PPDS?" edge cases (cake in a box, wrapped brownies on a stall, pre-orders); "may contain" for micro-producers (councils note blanket disclaimers are non-compliant, nobody explains the right way); compound/hidden ingredients ("what's in your sprinkles"); enforcement reality (a running fines/prosecutions tracker); batch-vs-recipe labelling.
4c. Lead magnets / free tools — exists vs missing
Exists: FSA's dry PPDS decision tool; static allergen-matrix templates everywhere; BakeProfit's no-signup calculators (their whole funnel); US nutrition-first label generators. Missing and winnable: an interactive scenario-based "Is my food PPDS?" checker; a live allergen matrix generator (type ingredients → matrix + hidden-allergen flags); a UK-native label preview generator (bolded allergens, first label free, no signup); a compound-ingredient decoder; a Bake Diary rescue page.
⚠️ Amended 24 Aug 2026 (AllergenKit deep dive): the free-tools gap is closing — AllergenKit now ships a live in-browser allergen matrix builder (click-to-tick, review-dated PDF, saves in-browser, no signup), a one-label-draft label maker and printable charts/posters, aimed at exactly our buyer; FoodCore's four free tools sit at the other end of the market. The pattern is now twice-validated AND first-mover on "live matrix, no signup" is partially taken — which raises the urgency of shortlist items 2 and 4 below and sharpens their spec: ours must generate the matrix from typed ingredients with hidden-allergen flags (the product demo in disguise), not just format hand-ticked boxes, which is all theirs does.
4d. Directories and the Bake Diary vacancy
Capterra/GetApp/Software Advice basic listings are free (paid PPC from ~$2/click with ~$500/mo floors — skip); G2 reportedly acquired Capterra from Gartner in early 2026 [extract-only — verify], so one profile strategy may now cover both. Serchen UK and SourceForge are cheap incremental listings whose category pages themselves rank for "bakery software UK". FoodCore is on all the majors; ProvenBatch is currently invisible. Week-one job at site launch.
Bake Diary (closed 31 May 2025, €6.95/mo, mostly UK/Irish bakers, no migration path): the "Bake Diary alternative" SERP is still being freshly contested with 2026-dated posts, and every recommended replacement is US/inventory-first — nobody mentions Natasha's Law/PPDS, which is exactly what the orphaned UK/Irish users need. A /compare/bake-diary page + "Bake Diary alternative UK" post is the fastest revenue per word written.
Reddit is the most-cited source across ChatGPT/Gemini/Perplexity/AI Overviews (🔻 noted, not actionable: Reddit is dropped permanently, 15 Aug 2026); review platforms (G2/Capterra/Trustpilot) carry a ~3× citation multiplier in recommendation answers; third-party-style comparison content is cited ~3× more than us-vs-them pages; original statistics lift citation rates 30–40%. AI referrals are ~1% of traffic but the fastest-growing and convert ~6× better than organic. AI Overviews are squeezing informational CTR (top-1 CTR down ~58% where they appear), so compliance-question posts must be built to be cited (schema, direct answers, original data) while commercial/tool pages carry the revenue.
4f. Realities for a new provenbatch.co.uk + prioritised shortlist
New domains: 3–6 months of suppression; only 1.74% of new pages reach top-10 within a year (Ahrefs via seo.ai); realistic meaningful rankings 6–12 months. This niche is moderately contested — FoodCore is real but young. Expect near-zero Google traffic before month 4; directories and the Bake Diary tail are the realistic signup sources in months 1–3 (Reddit answers — channel dropped permanently 15 Aug 2026).
Top content in priority order (full 15-item table in the research transcript; the load-bearing top eight):
"Bake Diary has closed — the UK alternative that does your PPDS labels too" + /compare/bake-diary
Interactive "Is my food PPDS?" checker (scenario-based, free, no signup — earns council/blog links)
Allergen labels for home bakers & cake makers: the complete UK guide (pillar; no software vendor ranks)
Free allergen matrix generator + supporting template post (the product demo in disguise)
ProvenBatch vs FoodCore, honest comparison (undercuts £19/£55 with £9; comparison pages get ~3× AI citations)
Why your label should come from the batch, not the recipe (differentiator manifesto — nearly empty space)
Best food labelling software UK 2026, including our competitors (contest FoodCore's double-dip)
"May contain": precautionary allergen labelling for small food businesses (identified gap; also our own open PAL question, RESEARCH.md §4)
hidden-ingredient decoder · "PPDS labels without a £300 printer" · wedding cakes · a Natasha's Law
fines tracker (verify SERP first) · a UK recipe cost calculator tying cost to label. Parallel week-one actions: free listings on Capterra UK, GetApp, G2, Serchen, SourceForge + genuine forum participation (Mumsnet, the Facebook groups — Reddit dropped permanently 15 Aug 2026).
5. Paid acquisition — the verdict is mostly no
5a. Why cold paid fails at £9–39/mo
At £14–15 blended ARPU and 4–5%/mo churn, gross LTV ≈ £175–280, so allowable CAC is £30–100 (3:1). Industry SMB-segment CAC benchmarks run $200–700 — 2–7× our ceiling; sub-$5k-ACV survivors do it with high-volume low-touch PLG, and "outbound economics rarely work" at this ACV. Funnel check: Meta's ~£22 blended cost-per-lead needs ≥22% lead→paid conversion to clear £100 CAC — not plausible from cold traffic. This is the standard "why paid fails at low ACV" result, confirmed across sources.
5b. Channel notes
Meta: the "small business owner" interest-targeting era is over (interests consolidated Jun 2025, discontinued Jan 2026; detailed targeting now treated as suggestions) — for a niche B2B offer the creative itself must filter ("If you sell cakes at markets…"). Lookalikes from customer lists survive and are the tool indie reports say works. New: Meta charges a 2% UK "location fee" on ads since 1 Jul 2026, billed outside Ads Manager metrics. Benchmarks: CPC ~£0.80–1.60, CPL £20–40 for niche B2B. Realistic test: £300–600/mo for 6–8 weeks minimum — defer until organic proves messages.
Google Search: UK all-industry CPC ~£1.95; B2B SaaS £3–9. No published CPC exists for "allergen label software"/"natasha's law" — almost certainly low-volume, low-competition long tail (their SERPs are SEO content, not ads). The one cold paid bet worth making: exact-match Search at £5–10/day. Manual check for Dave first (dashboards, no CLI): Google Keyword Planner volumes + adstransparency.google.com and the Meta Ad Library for nutritics.com, kafoodle.com, foodcore.io, planglow.com, brother.co.uk — both 403'd from this session.
Performance Max: avoid (needs ~30 conversions/mo to learn, treats spam form-fills as wins, cannibalises Search). LinkedIn: no (economics built for $25k deals; audience barely there). TikTok ads: plausible later, organic first (B2B CPL $25–60 reported, 30–50% below LinkedIn — agency-sourced). Pinterest: cheap clicks (~$0.30–0.83), wrong mode (recipe-browsing, not compliance) — at most a £5/day promoted-template test.
Where paid does fit: (1) exact-match Search as above; (2) retargeting + boosting proven organic posts (£20–50/boost) once the site has traffic; (3) customer-list lookalikes after ~100 customers. Never: LinkedIn, PMax, broad cold Meta prospecting.
6. Email, automation, and the legal rails
ProvenBatch is a compliance product; its own marketing must be visibly compliant. This section separates LAW from TACTICS. (Direct ICO pages 403'd; every legal point below is corroborated across ≥2 independent legal/ICO-derived extracts. The ICO's updated direct-marketing guidance was due Spring 2026 — re-check before the first campaign.)
6a. LAW — marketing email (PECR + UK GDPR + DUAA)
Consent is the default for marketing email to individual subscribers. No small-business or "it's B2B" exemption.
🔴 Sole traders and non-LLP partnerships (E/W/NI) are INDIVIDUAL subscribers under PECR. Most of our audience are sole traders, so cold email to them without consent/soft opt-in is a PECR breach — "B2B cold email is fine" is false for this market. Cold email to limited companies is lawful (identify yourself, honour opt-outs, LIA under GDPR) — which makes cold outreach a manual, Companies-House-checked, low-volume tactic for delis/farm shops at most, not a channel.
Soft opt-in, exactly (all three, cumulatively): details obtained in the course of a sale or negotiations for a sale; marketing only our own similar products; free opt-out offered at collection and in every message. A trial signup plausibly qualifies as "negotiations for a sale" (practitioner interpretation, not an ICO ruling); a lead-magnet download alone does NOT — put an unticked "send me PPDS tips" checkbox on every magnet form and log it.
DUAA changes (in force 5 Feb 2026): soft opt-in extended to charities (irrelevant to us); the commercial rules and sole-trader status are unchanged; maximum PECR fines rose from £500k to £17.5m / 4% of turnover. Enforcement lands on ordinary-sized firms (HelloFresh £140k; Allay £120k + ZMLUK £105k in Jan 2026; average PECR fine ~£95k across 49 actions since Mar 2022), and the near-universal root cause is inability to prove consent existed at send time → keep a consent log (timestamp, source, wording) — a trivial Supabase table.
6b. LAW — cookies and tracking
DUAA created new consent exceptions to PECR reg 6 (5 Feb 2026): first-party statistical/analytics cookies used solely to improve your own service are now banner-free — but the exception is lost if data feeds advertising or profiling (so GA4 is doubtful), even exempt cookies need clear information + a simple opt-out, and Meta/Google ad pixels still require prior opt-in consent. Cookie breaches now carry the £17.5m/4% maximum.
TACTICS — ship banner-free: Plausible or Fathom (cookieless, ~£7–9/mo, no banner by design); Cloudflare Web Analytics has conflicting evidence on whether it sets cookies [unresolved — test what it actually sets, since we're on Cloudflare anyway]. No ad pixels initially — a compliance product whose site opens with a grudging tracking banner undermines the pitch. Attribute via UTMs + a "how did you hear about us?" signup field + per-channel promo codes.
6c. TACTICS — capture, nurture, lifecycle
Lead magnets matched to compliance fear: PPDS checklist, 14-allergens poster, label template pack, "Does PPDS apply to me?" quiz, EHO-visit prep checklist. Benchmarks (directional): checklists/templates 20–40% on dedicated pages; quizzes ~40%; one form field beats two (4.4% vs 2.9%); single-CTA pages 13.5% vs 2.5% with 5+.
Tooling at near-zero cost: keep Resend for transactional (already the stack, D-09); MailerLite free tier (~1,000 contacts, automations + landing pages) or Loops free tier (1,000 contacts, event-driven) for marketing/lifecycle; £0/month to start. Double opt-in ON — not legally required, but it is the consent proof the ICO fines people for lacking.
Trial lifecycle (5 emails, one behavioural trigger): day-0 welcome with a single activation goal ("print your first compliant label"); day-2/3 nudge only if no label created (targets the documented 72-hour activation cliff — 68% of trials not activated in 72h never convert); day-7 proof/objection (an EHO-inspection story); trial-end-minus-2 with clear pricing; day+3 win-back. Wire off a label_created event. Promotional sends during trial ride the soft opt-in — state it at signup and carry the unsubscribe everywhere.
Pre-GA list building: value-first content + waitlist 3–6 months ahead (i.e. now — GA is 19 Jan 2027), fed by the community/content work in §2–4, not by paid spend.
Addendum, 30 Aug 2026 (RESEARCH.md §19d has the full pass). The list-building half is now decided: near-term the list lives on our own rails — /beta's updates_opt_in consent into beta_requests, updates-only capture + unsubscribe + campaign sending filed as issues from the content-upgrade session — because a pre-launch list does not qualify for PECR's soft opt-in and our unticked checkbox already collects the explicit consent correctly. When the list outgrows the rails (~250+ subscribers or GA), MailerLite is the pick over Loops: EU data residency + a DPA on every account by default, built-in double opt-in and GDPR forms, free to 250 subscribers/2,500 emails a month, then from $12/mo (its free tier is smaller than the ~1,000 contacts stated above — that figure has aged). Buttondown (API-first, 100 free, ~$9/mo) is the runner-up; Kit cliffs to $39/mo; Mailchimp's free tier is no longer competitive. Adopting any of them is a new sub-processor → a D-12 disclosure change, priced into the decision.
7. Events — visit, don't exhibit
7a. The exhibiting arithmetic
UK shell scheme runs £220–625/sqm; realistic all-in £3,000–5,000 per national show. Trade-show ROI literature is blunt: shows pay off where deal size absorbs the cost, and are explicitly "not worth it for low-value, self-serve products". At ~£300 LTV a £4k stand must convert 13+ paying customers from the stand alone — implausible for v1. 2026–27 answer: visit everything (trade tickets are free-to-cheap), exhibit nowhere national. The only exhibit experiments worth modelling later: a regional show (The Source, £/sqm at the bottom of the range) or FRA conference sponsorship.
7b. Calendar (dates re-verified where flagged; all extract-only)
Event
Dates
Venue
Why
Speciality & Fine Food Fair
8–9 Sep 2026 ⚠️ (official socials; older listings say 15–16 — re-verify before travel)
Olympia
~700 exhibitors, 10k trade attendees; artisan producers/delis; FRA is an official partner — two channels, one visit, pre-beta
lunch! (+ Artisan Food & Drink Show)
16–17 Sep 2026
ExCeL
Food-to-go and cafés = PPDS ground zero
Street Food Business Expo / Food Service Industry Expo (7 co-located shows)
Free ticket; traders + takeaways + cafés; NCASS runs the Keynote Theatre — open that conversation in person, pre-beta
Cake International / Bake International
6–8 Nov 2026
NEC Birmingham
The densest gathering of home/business cake makers in the UK, in our beta month, an hour from Aylesbury. Go with beta invite cards; audit whether any software exhibits. Exhibitor pricing unpublished (contact: melanieu@ichf.co.uk)
The Cake & Bake Show
26–29 Nov 2026
Olympia
Consumer-skewed; only if London-convenient
The Source Trade Show
2–3 Feb 2027
Westpoint, Exeter
First show after GA; regional, ICP-dense; the venue to price a future exhibit test
Northern Restaurant & Bar + lunch! NORTH
9–10 Mar 2027
Manchester Central
Pair with northern customer visits
Farm Retail Association Conference
~Mar 2027, TBA
TBA
Few hundred farm-shop decision-makers; may be the one event where sponsoring beats visiting — chase dates
UK Food & Drink Shows (Farm Shop & Deli Show + Food & Drink Expo + 2 more)
12–14 Apr 2027
NEC
25k visitors, four shows, one free ticket, post-GA
NMTF Young Traders Markets / Bakers' Fair
rolling / unverified
various
NMTF is a partnership, not an event visit; Bakers' Fair 2026 dates need bakeryinfo.co.uk from an unblocked connection
Local food festivals & makers markets (Thame Food Festival ~late Sep etc.)
rolling
local
Best guerrilla channel: walk stalls as a customer and talk PPDS — the traders are behind the stalls. Visit, never exhibit
8. Partnerships & referral channels
Ranked, with the evidence:
Food hygiene training providers — High Speed Training first. The market leader in Level 2 e-learning runs a public affiliate programme on Paid On Results (rate not visible), and its Food Hygiene Hub ranks for exactly the searches new food businesses make. Pitch a tools mention/co-content now, affiliate mechanics later. Virtual College publishes a "starting a food business from home" guide — same play. The rails already exist; this is the lowest-effort route.
Local authority / EHO signposting, starting with Buckinghamshire. Every customer legally passes through a council page 28 days before trading. Councils overwhelmingly link the FSA, but demonstrably host audience-specific guidance when it exists (West Norfolk's "cakemakers guidance" PDF). The ask is signposting alongside FSA links, and a free tier materially helps. Expect "we can't endorse commercial products" by default; start with one friendly EHO conversation. [KNOWLEDGE-ONLY, verify: the FSA's Safer Food Better Business packs have no software companion — a "digital sidecar to your SFBB pack" positioning is untested but unclaimed.]
NMTF before NCASS. Both own their verticals (25k traders; 5k+ caterer members). NCASS has channel-conflict risk (bundles its own compliance offering into membership) — approach as "complements membership", and lead with NMTF, where no competing product surfaced. A member-benefit listing or Young Traders Market sponsorship is the cheapest stallholder route.
Non-software label printers & food wholesalers. Planglow proves labels+software bundles sell (they are a competitor, not a partner — and their compliance-fear content marketing validates ours). Pilgrim Foodservice already markets a "one-stop Natasha's Law compliance solution" to its wholesale customers — the proof that wholesalers will carry this message; BFP/Bako-type bakery wholesalers are the analogous targets [unverified]. Plain label printers (e.g. Reflex Labels — PPDS content, print only) are the clean non-competing bundle: "we generate the label, they print the stock".
Insurers (CMTIA first). The compliance-condition link is real and sourced: Protectivity's home-bakery liability pays out "provided you have followed safety and hygiene standards" — "your cover assumes you're compliant; here's the tool" is a strong story. CMTIA (market-trader policies £65–110/yr, small Bristol broker) is reachable founder-to-founder for a discount-code swap; Simply Business's content hub is the SEO prize. No precedent of an insurer co-marketing a compliance SaaS was found — treat as experiment, not plan.
Open gaps (search budget): accountants/bookkeepers for food micro-businesses, kitchen-share operators, SFBB distribution detail, affiliate commission rates.
9. The channel plan that falls out, mapped to the launch timeline
Phasing against launch-plan.md (beta invites 23 Nov 2026 · first revenue 11 Dec 2026 · GA 19 Jan 2027):
When
Channel work
Cost
Aug–Sep 2026 (M1/M2)
Join the FB groups and participate (D-17 already says so); start the TikTok/Reels asset drumbeat when #172's videos exist; visit SFFF (8–9 Sep) and Street Food Business Expo (29–30 Sep) — open the NCASS conversation in person; start the waitlist + first lead magnet; directory profiles can't be claimed until the site exists — prep the copy
£0 + train tickets
Oct–Nov 2026 (M3, site launches)
Week-one directory listings (Capterra/GetApp/G2/Serchen/SourceForge); publish content #1–4 (Bake Diary pages, PPDS checker, home-baker pillar, matrix generator); walk Cake International 6–8 Nov with beta invite cards; first admin-partnership approaches (Bennett, A Bigger Slice, Tynan); first HST/council conversations
£0–£200
Dec 2026–Jan 2027 (M4→GA)
Trial lifecycle emails live (soft opt-in stated at signup, consent log on); nano-creator experiments (2–3 coaches, £50–150 each) + one £50 OLBAA podcast slot; exact-match Google Search test at £5–10/day once Keyword Planner confirms volume; content #5–8
£200–£500/mo
Post-GA 2027
Ambassador/affiliate programme (recurring commission, CakeFlix/Craftybase as templates); NMTF member-benefit deal; insurer/wholesaler experiments; The Source (Feb) and UK Food & Drink Shows (Apr) visits; scale what converted, kill what didn't — the "how did you hear about us?" field is the arbiter
scales with revenue
The one-line strategy: organic and partnerships carry acquisition (communities via admin partnerships, SEO into the vertical gaps FoodCore ignores, nano-creators on commission, councils and trainers as referrers); paid is a £5–10/day exact-match Search experiment and nothing else until ~100 customers exist; and every email practice is consent-logged because our marketing is part of the product's compliance story.
10. Follow-ups that would harden this most
Read the 2019 PPDS Impact Assessment (legislation.gov.uk/ukia/2019/144) from a browser — the official in-scope business count is the single highest-value missing number (drives TAM).
Query the FHRS open-data API (api.ratings.food.gov.uk, business-type field) from an unproxied machine — real segment counts including mobile caterers and home caterers; or FOI the FSA for domestic-premises totals (drives SAM).
30 minutes logged into Facebook: verify member counts and promo rules for the §2 shortlist.
Google Keyword Planner + Ads Transparency Center + Meta Ad Library (all dashboard-only from here): volumes for the §4 keyword set; whether FoodCore/Nutritics/Kafoodle/Planglow run ads.
Chase FRA 2027 conference dates and Bakers' Fair 2026 dates; re-verify SFFF (8–9 vs 15–16 Sep) before booking travel.
Re-check the ICO's updated direct-marketing guidance (was due Spring 2026) before the first email campaign.
Market
Last updated 20 Sep 2026
31. Where the customers are and how to reach them — acquisition research for the 1 Oct public trial (researched 20 Sep 2026)
Dave's brief: two questions — where are PPDS-obligated small food businesses, and how do we find them lawfully at near-zero cost; and how do we reach them so they start a trial and pay. Nine parallel passes, every claim graded verified / reported / inference, dated and linked. Full write-up: outputs/docs/gtm-acquisition-research-2026-09.md (§C11 is the one-page 90-day plan; §C12 the twenty named leads). Builds on §12 / gtm-channels-research.md and challenges it only where this pass found better evidence:
The home-based stock has a primary-data anchor. FHRS open data marks a private-address business by a blank address and an outward-only postcode (GOV.UK, 30 Jun 2026). A 14-authority sample (28,163 records, pulled 20 Sep) puts that pattern at 16% of all records and 62–63% of "Other catering premises" / "Mobile caterer"; scaled to the 612,609-record register that is ~95–100k establishments at a private address (±20%) — inside §12's 60k–120k model. The AwaitingInspection rating is in practice the public new-registration feed (9% of records, 18–23% in home-type categories). FHRS holds no email or phone, so it is a targeting and measurement source, not an outreach list.
FSA Sept 2026 Board: 104,000 new registrations in 2025/26 (+8.7%), ~590k establishments, 36,000 awaiting first inspection. §12 used ~90k/yr.
Three corrections to what §12 assumed. AllergenKit is live at £9/mo for "cake sheds · home bakeries · market bakers" — the sub-£10 slot is not empty (the 24 Aug deep dive already knew). Pilgrim Foodservice's "Natasha's Law solution" is Planglow's LabelLogic Live via Erudus — an occupied wholesaler rail; the realistic wholesaler ask is an Erudus data integration. High Speed Training's affiliate programme pays outward (20–30% to sites that send HST customers); it is not HST referring to us. Councils and the FSA link no third-party tools (ten pages, zero links).
The per-trial CAC ceiling does not close. "£6–20 per started trial" only reconciles with the £30–100 allowable CAC at ~20% trial-to-paid; the best opt-in no-card benchmark (ChartMogul with Kyle Poyar, Jan 2026, 200 products, US/global) is 4–6% good, 10–15% great, median 8%. Either activation earns the "great" band or the paid ceiling per trial is ~£1.50–£8.
ICO direct-marketing guidance is current for the DUAA (updated 28 Apr 2026) and says in its own words that "signing up to a free trial" is "negotiations for the sale" — the soft opt-in covers trial signups if an opt-out is offered at signup and in every message. PECR fines rose to £17.5m/4% on 5 Feb 2026 (SI 2026/82). Recognised legitimate interests do not include direct marketing. Lawful first contact without consent: sole trader — TPS-screened call, letter, in person; Ltd/LLP — the same plus email with identity and opt-out.
Message. "Fear > time > cost" is half right: fear triggers first-timers; established hand-labellers' live pain is keeping labels correct when packs and recipes change; cost is a barrier to software, not a motive. The missing angle is customer trust — FSS 2023: allergic consumers avoid small outlets ("I would like to buy more locally, but I don't trust them"); FSA Our Food 2023: 17 of 47 PPDS samples had an undeclared allergen, failures "restricted to smaller food businesses" and mostly no label or no ingredients list rather than a wrong list. Never claim hand-written labels are illegal (GOV.UK says they are permitted if legible).
Referral: ship "give a month, get a month" first (FreeAgent's pattern — vests on the referee's first payment, never during a trial, stacks to a free year), shown on the read-only banner; no cash, no affiliate software (all ≥$39–49/mo). Coach affiliates post-GA at the 20–30%-for-12-months norm (Stocksmith, Xero, CakeFlix).
Bake Diary's closure is the loudest objection in the buyer's own groups (two weeks' notice, users lost data). Recommendation: publish a wind-down commitment (notice period, export format) on the pricing page.
Dead leads: Guild of Jam and Preserve Makers (domain redirects elsewhere), Karma Kitchen (domain for sale). Contested SEO gap: BakerInbox (UK) took "Bake Diary alternative UK" on 15 Jun 2026.
Market
Last updated 20 Sep 2026
Where the customers are and how to reach them — evidence-graded acquisition research for the 1 October 2026 public trial
Researched 20 Sep 2026, nine parallel research passes (where businesses are visible · trigger moments · referrers · analogues · message evidence · lawful outreach · trial-to-paid · referral and objections · named leads), synthesised into one document. Companion to gtm-channels-research.md (31 Jul 2026 — the channel map this builds on and does not repeat), metricool-growth-research.md (15 Sep 2026 — the paid-and-social loop) and go-to-market-decisions.md (what is decided). This document answers two questions: where are our customers and how do we find them lawfully at near-zero cost; and how do we reach them so they start a trial and then pay. Everything already known from the July pass is taken as read and only challenged where this pass found better evidence — each challenge is flagged in §0.
How to read the grades. Every material claim carries one of three grades, next to the claim: [verified] = the primary page was fetched and read (or the API/data file was pulled and parsed) on 20 Sep 2026; [reported] = a secondary source or a search extract that was not opened; [inference] = our reasoning from the evidence. [repo] marks a fact taken from a ProvenBatch repository document or the production database — primary for us, but internal. Where something could not be checked it says "not verified" rather than guessing. Dates are the date on the page where one exists, otherwise "undated".
Two limits on this pass, stated once. (1) The session's shared web-search budget ran out part-way through every pass; the second half of each was done by fetching known primary URLs directly, so coverage is deep on official sources and thinner on forum and social quotes. (2) The outbound proxy returned 403/405 on a set of sites (Pilgrim Foodservice, Avery, Brother's blog, British Baker, Etsy's policy pages, the FSA's own alerts site, TPS Online, Meta's Ad Library, several council PDFs); each is named where it matters and listed at the end.
0. Headlines — what this pass changes, and where it contradicts §3 of the brief
The home-based stock now has a primary-data anchor. The FHRS open data marks a business trading from a private address by publishing a blank address and an outward-only postcode (GOV.UK guidance updated 30 Jun 2026 [verified]). Applying that test to 14 random authorities (28,163 records, pulled 20 Sep 2026) gives 16% of all records and 62–63% of "Other catering premises" and "Mobile caterer" records [verified, own analysis]; scaled to the 612,609-record register that is roughly 95,000–100,000 UK establishments at a private address [inference, ±20%]. It sits inside the brief's 60k–120k model and should be recorded as its first primary-source check. §A1.
New-registration volume is higher than the July figure. The FSA's September 2026 Board paper reports 104,000 new registrations in 2025/26 (+8.7%), ~590,000 establishments and 36,000 premises still awaiting a first inspection [verified]. The July pass used ~90k/yr. §A2.
🔴 Contradiction: the sub-£10 slot is not empty. AllergenKit (allergenkit.co.uk) is live at £9/month, 7-day no-card trial, founding-member price lock, aimed at "cake sheds · home bakeries · market bakers", with free no-signup tools and a concierge "tell Finn one recurring bake" offer [verified]. The repo's own 24 Aug deep dive already records this; the brief's line is stale and should be corrected. §A4.
🔴 Contradiction: Pilgrim Foodservice's "Natasha's Law solution" is Planglow's LabelLogic Live fed by Erudus (Planglow, 7 Dec 2021 [verified]; price £120/yr [reported]). It is an occupied wholesaler rail, not an open one. Booker (Apr 2024) and Bidfood (DayMark plus its own MyRecipes) have likewise chosen. The realistic wholesaler ask is an Erudus data integration, not a co-marketing deal. §A3, §A4.
**🔴 Contradiction: High Speed Training's affiliate programme pays *outward*** — 20–30% of order value to sites that send HST customers (Paid On Results [verified]). It is a way for us to earn a few pounds per Level 2 referral, not a way for HST to refer to us. Virtual College (Awin, 10%/5%) is the same shape. §A3.
Councils and the FSA do not link third-party tools. Ten FSA, council, growth-hub and CIEH pages read; zero third-party links [verified]. The July pass's "councils host third-party guidance" rail rests on one West Norfolk PDF; treat EHO teams as an audience to make the product legible to, not a referrer. §A3.
The trial-to-paid arithmetic in the brief does not close. Allowable CAC £30–100 and "£6–20 per started trial" reconcile only at ~20% trial-to-paid. The best primary benchmark for an opt-in, no-card trial (ChartMogul with Kyle Poyar, Jan 2026, 200 self-serve products) is 4–6% "good", 10–15% "great", median 8% [verified; US/global]. At 4–8%, £6–20 per trial is £75–500 per customer. Either the activation work in §B8 earns the "great" band or the paid ceiling per trial must fall to about £1.50–£8. §B8.
The ICO's direct-marketing guidance is now current for the DUAA (updated 28 Apr 2026 [verified]) and its own wording says "signing up to a free trial of your product or service" is "negotiations for the sale" — the soft opt-in is available to us for trial signups, provided an opt-out is offered at signup and in every message [verified]. The Guide to PECR pages are still marked "under review" and the "PECR advice for small organisations" update was listed as Drafting/Summer 2026 with no page found — pending. §B7.
"Bake Diary alternative UK" is contested, not open. BakerInbox (UK, £19/mo) published the UK page on 15 Jun 2026; four non-UK sites hold it too; none of them does Natasha's Law labels [verified]. Still worth writing, but it is a differentiated page now, not a vacant one. §A4.
Nothing found supports contacting an individual business at a trigger moment. Every trigger with real evidence (registration, calendar, enforcement stories, recalls, marketplace gates) is reachable through a public page, a search result or an admin-approved group post, never a list. The lawful first-contact matrix (§B7) is the same for every source of leads: post, in person, TPS/CTPS-screened calls, corporate email to Ltd companies only, and public replies where the business asked in public.
Two named leads in the brief are dead. The Guild of Jam and Preserve Makers domain now redirects to an unrelated site; Karma Kitchen's domain serves a for-sale page [verified]. Both are excluded from §C12.
A. Finding customers
A1. Where PPDS-obligated small food businesses become visible
Bottom line. The one lawful, free, national, daily-refreshed place every vertical becomes visible is the FSA's FHRS open data (612,609 establishments, OGL v3, no API key), and it carries two usable signals: a business type, and the private-address pattern that marks a home business, plus an "AwaitingInspection" rating that is in practice the public new-registration feed [verified]. It holds no email or phone, so what it lawfully enables is targeting and measurement, letters to shop addresses, and in-person visits — not electronic outreach; every other source (council registers, market directories, association finders, platforms) is small, patchy or blocked.
A1.1 Council registration and the FHRS open data
What it is. The FSA publishes hygiene ratings for every local authority as per-LA XML/JSON files and a REST API; "updated daily"; the full dataset held 612,609 businesses on 17 Sep 2026 [verified, https://ratings.food.gov.uk/open-data]. "No sign-up process, API keys, or login details are required" [verified, https://api.ratings.food.gov.uk/help, undated]. Licence: Open Government Licence v3, update frequency "near real-time" [verified, https://data.food.gov.uk/catalog/datasets/38dd8d6a-5ab1-4f50-b753-ab33288e3200]. OGL lets you copy, adapt and "exploit the Information commercially" with attribution, and excludes personal data [verified, nationalarchives.gov.uk/doc/open-government-licence/version/3/]. Cost: free.
Fields per record (parsed from FHRS417en-GB.xml, North Tyneside, extract 16 Sep 2026) [verified]: FHRSID, LocalAuthorityBusinessID, BusinessName, BusinessType, BusinessTypeID, AddressLine1–4, PostCode, RatingValue, RatingKey, RatingDate, LocalAuthorityCode/Name/WebSite/EmailAddress, Scores (Hygiene, Structural, ConfidenceInManagement), SchemeType, NewRatingPending, Geocode. The API adds Phone (empty in every record inspected), RightToReply, Distance [verified, live call 20 Sep 2026]. There is no business email field; the only email is the council's environmental-health inbox. RatingValue includes AwaitingInspection, AwaitingPublication, Exempt and (FHIS) Pass / Improvement Required [verified].
Business types (live from /BusinessTypes/basic, 20 Sep 2026) and how our verticals land [verified counts; mapping is inference from names seen in the files]:
BusinessTypeId
Name
UK count
Our verticals that land here (examples seen in the files)
7841
Other catering premises
76,033
Home bakers ("Biddys Bakes", "Whisked by Caitlin"), meal-prep ("shredosmealprep"), granola/preserve makers
7846
Mobile caterer
31,317
Street food and stalls ("Hot Wheels Street Food", "The Loaded Trailer")
4613
Retailers - other
120,326
Some home bakers, butchers ("Charnwood Wild Venison & Game"), delis, farm shops, honey/preserves
7839
Manufacturers/packers
12,138
Chocolatiers, preserve and sauce makers, some cottage bakeries
1
Restaurant/Cafe/Canteen
141,010
Cafes packing sandwiches
7844
Takeaway/sandwich shop
63,041
Sandwich shops
7838
Farmers/growers
2,907
On-farm shops
-1
All
612,609
Caveat: type is assigned by the LA and is inconsistent — the same kind of home cake business is 7841 in one authority and 4613 in another [inference from samples]. Any pipeline filters on all of 7841, 7846, 4613 and 7839 (plus 1 and 7844 for cafes) and then classifies by name.
The home-based signal exists and is strong. GOV.UK: "If your business is registered at a private address (for example you are a home caterer), only the first part of the postcode is published… You may give permission for the full address to be published" (guidance updated 30 Jun 2026) [verified, gov.uk/government/publications/food-hygiene-rating-scheme-fhrs-guidance-for-businesses]. So the machine-readable test is address lines empty and postcode outward-only. Applied to 14 random FHRS authorities (Maldon, Darlington, North Yorkshire, Blaenau Gwent, Bristol, Swansea, Cambridge, Newham, Walsall, Bracknell Forest, Spelthorne, Folkestone & Hythe, Bexley, Burnley; 28,163 records; files pulled 20 Sep 2026) [verified, own analysis]:
BusinessType
Records
Private-address pattern
AwaitingInspection
Awaiting AND private
Other catering premises
2,941
1,811 (62%)
527 (17.9%)
411
Mobile caterer
1,430
903 (63%)
335 (23.4%)
233
Manufacturers/packers
352
103 (29%)
55 (15.6%)
31
Retailers - other
6,187
1,007 (16%)
415 (6.7%)
199
Farmers/growers
117
90 (77%)
27 (23.1%)
25
Restaurant/Cafe/Canteen
6,898
316 (5%)
589 (8.5%)
140
Takeaway/sandwich shop
2,827
78 (3%)
240 (8.5%)
15
All types
28,163
4,516 (16%)
2,528 (9.0%)
1,094
Per-LA private share ran from ~11% (Newham, Burnley) to ~25% (Maldon, Bracknell Forest). [inference] Scaling 16% to 612,609 gives ~95,000–100,000 establishments trading from a private address, of which ~45,000–50,000 are "Other catering premises" or "Mobile caterer". The pattern also catches some non-PPDS records (childminders, distributors, allotment egg sellers) and misses home businesses that consented to full-address publication — treat as ±20%. This is the first primary-data anchor for the brief's 60k–120k modelled range; report it as such, not as a replacement.
New registrations. The FSA publishes no new-registration list; the register.food.gov.uk privacy notice says the FSA passes details to the LA and "the list of registered food business establishments is published by the relevant local authority", citing Article 113 of retained Regulation (EU) 2017/625 [verified, https://register.food.gov.uk/privacy-notice, undated]. In practice the FHRS AwaitingInspection record is the public new-registration signal: 9.0% of all sampled records, 18–23% in the home-type categories, with a blank RatingDate [verified]. In Charnwood (87 awaiting, 20 Sep 2026) the list read like our customer list — "Elsie's Cakery", "Jessica Blossom Bakes", "Cookie For Thought", "Nom Granola", "shredosmealprep", "Hot Fresh Donuts" — all private-address [verified]. Two API findings: the ratingKey filter on /Establishments did not filter in tests (every variant returned the full count) and there is no sort-by-date-added; to detect new registrations you diff the daily per-LA files on FHRSID (~360 files, a cheap nightly job) [verified].
Can we use it lawfully for outreach? OGL does not licence the personal data inside the dataset, and a sole trader's business name is usually personal data — "If you can identify an individual either directly or indirectly it will constitute personal data even if they are acting in their business capacity" [verified, ICO B2B marketing page]. So processing rests on our own lawful basis (legitimate interests for postal B2B marketing, with the right to object honoured) [inference from the ICO page]. Under PECR, sole traders and ordinary partnerships are individual subscribers: no email, no SMS, and — because the ICO's definition of electronic mail covers "direct messaging on social media" — no social DMs without consent or soft opt-in [verified, ICO electronic-mail guidance]. Corporate subscribers (Ltd, LLP, Scottish partnership) may be emailed with identity disclosed and an opt-out [verified]. Live calls must be screened against both TPS and CTPS [verified]. Post is outside PECR [verified]. Therefore the lawful first contacts an FHRS lead enables are: (a) a letter to a published trading address (shops and units, not home businesses, whose address is withheld); (b) in person at the market or shop; (c) a TPS/CTPS-screened call — but FHRS holds no number; (d) corporate email only after a Companies House match, which most home bakers fail; (e) a public reply where the business asked in public. Full matrix in §B7.
What FHRS is best for [inference]: not outreach but targeting and measurement — choosing LAs and postcode districts for the £150/mo paid plan, counting the home-based stock, seeding per-town SEO pages, and knowing which councils hold the most awaiting-inspection home businesses (the EHO audience). TPS/CTPS costs are reported only (direct licences "from £300" TPS / "from £150" CTPS; third-party checks from £4.25+VAT per 325; Selectabase, Aug 2026 [reported — the TPS site itself did not render]).
A1.2 Council public registers
GOV.UK (updated 25 Jun 2026): anyone who sells, cooks, stores, handles, prepares or distributes food must register with their LA at least 28 days before trading; home-based, mobile and online sellers included [verified, gov.uk/food-business-registration]. Council pages still describe a public register open to inspection — Salford: "A register of addresses and the type of business carried on at each will be open to inspection by the general public" [verified]; Ealing lists name, address, telephone and type "available for inspection by the public" [verified] — but inspection is at council offices and there is no national machine-readable version. A few publish online: North Norfolk (PDF) [verified page; PDF not opened]; Cornwall's "Food Register" dataset on data.gov.uk, OGL, last updated 29 Sep 2020 [verified]. FOI precedents exist (Enfield 2012, B&NES, Westminster on WhatDoTheyKnow) [reported — 403]; councils routinely withhold home addresses under s.40 [inference]. No council publishing a "new registrations this month" list was found. Cost: free. Lawful use: same as FHRS — a shop address at best.
A1.3 Market operators, farmers'-market and food-hall directories
Place
What is published
Scale, freshness
Lawful use
Cost
Farm Retail Association "Find a Farm Retailer" (farmretail.co.uk/find-a-farm-retailer/)
Name, description, category, contact on the member page; map search; public
No public trader directory; "Market Near Me" (522 on 20 Sep 2026); affiliates include Pedddle [verified]
~25,000 members stated [verified]
Partner deal or a piece in member comms
£150/yr sole-trader membership [verified]
Pedddle (pedddle.com/stalls/)
Public stallholder profiles by category and county
7 Food & Drink stallholders vs 137 Clothing & Jewellery; 2,136 markets listed in the past year [verified]
Self-published → public reply or visit; no DMs
Listing £60/yr or £15/mo [verified]
farmers-market.org.uk
294 markets, 48 organisers, 102 stallholders with category (Bakery, Butcher/Meat, Honey & Preserves, Sweets & Chocolate, Hot Food) and markets attended [verified]
Current but tiny
Visit / public reply
Free to read
NCASS "Find a Caterer"
A request-a-quote form needing an account — not a directory [verified]; "30+ partner deals" and an Allergen Hub [verified]
Membership count not stated
Partner route only (channel conflict: NCASS bundles its own Digital Food Safety System [verified])
—
Nextdoor Business Pages (nextdoor.co.uk/directories/)
Free page; public directory with a "Bakery" category [verified]
Unknown
Business posted publicly → public reply lawful
Free
Council market pages, Bury Market list, KFMA, Love British Food, localproducer.co.uk
Blocked (403/DNS)
—
—
not verified
Butcher and fine-food bodies: National Craft Butchers' "Find a Craft Butcher" URL returned 404; Q Guild's member page 403; Guild of Fine Food 403 throughout — not verified.
A1.4 Local-authority "home baker" guidance pages
These are where a new home business reads about PPDS, so they matter for message and for placement, not as lists. North Yorkshire's Food safety pack for home bakers covers registration, SFBB, traceability and PPDS labelling with the 14 allergens in bold [verified, northyorks.gov.uk]; Scottish Borders' pack (25 Jun 2021) links the FSA's PPDS guidance [verified]; Reigate & Banstead and Wigan publish printable packs [verified exist]; King's Lynn & West Norfolk hosts four cake-maker PDFs (allergen guide, checklist, packaging and labelling, cakemaker's guidance) [verified]. None publishes who registered, and — see §A3 — none links a third-party tool.
A1.5 Training providers, insurers, wholesalers
Training. High Speed Training runs a named case-study series ("The Savvy Baker", 21 Aug 2026; GOAT; Sufra; Waterstones Café W; Taylors of Harrogate) [verified] — a handful of already scaling businesses, not a cohort list. Virtual College: no case-study page (404). NCASS sells Level 1–3 and allergen courses to members; cohorts not public [verified]. Coach graduate showcases: not checked (budget).
Insurers. Protectivity (knowledge centre, refer-a-friend £25/£25, no community) [verified]; CMTIA (newsletter and blog, no directory) [verified]; Superscript (newsletter; blog Jul–Sep 2026) [verified]; Simply Business (knowledge centre; a 2021 bakery feature now 404) [verified]; PolicyBee (no cake-maker page; food trades absent from its list) [verified]. No insurer publishes customer lists or a community; their blogs are a content-partnership surface (§A3).
Wholesalers. Bako launched an allergen-checker portal by product code before Natasha's Law (British Baker 2021) [reported — 405]; Pilgrim markets Planglow's LabelLogic Live [reported — 403; confirmed from Planglow's side, verified]; Booker is an NCASS partner deal [verified]; Brakes is an FRA partner [verified]; BFP rendered empty; Bidfood and Costco not checked. None publishes customers; the only visibility a wholesaler offers is physical — trade counters and cash-and-carry car parks, where in-person contact is lawful [inference].
A1.6 Cottage-food marketplaces and ordering platforms (UK reality check)
Platform
Real, UK?
Public directory?
Notes
Bakesy (bakesy.app)
Real; US-oriented; UK availability not stated [verified]
No directory; per-baker storefront; 30-day trial; no labelling feature mentioned [verified]
UK App Store 4.39 from 87 ratings [verified]
findyourbaker.com
Real directory
UK page lists 109 bakers with town, specialties, some phones [verified]
Real UK marketplace; "Sell with us" via BakerCentral [verified]
Listings exist; count not stated
—
Etsy
Food allowed subject to local law and labelling; perishables and unlabelled allergens prohibited [reported — policy page 403]
Search only; shop pages show name and town
Sole traders; "message seller" is for buying, not marketing
Nextdoor Business
UK; free pages; Bakery category [verified]
Yes, by place
See A1.3
Facebook Marketplace / Shops, Instagram Shops
Meta's commerce policy page says nothing on food [verified partial]
No seller directory
Scraping is off the table anyway
Shopify, Square, SumUp
No public seller directories verifiable
—
not verified
Cake Shed, Jam Jar, Homemade.io, Orderly, Shoppable
Not checked (budget) — do not cite either way
—
—
A1.7 Summary — how findable, how many, how fresh
Place
Findable
How many
Fresh
FHRS open data / API
Trivial, free, national, machine-readable
612,609 total; ~76k Other catering + 31k Mobile + 12k Manufacturers + 120k Retailers-other; est. ~95–100k private-address; est. ~55k awaiting inspection [inference from 9%]
Daily
Council registers
Poor: office inspection, a few PDFs, FOI
Per LA
Stale (Cornwall 2020)
FRA finder
Easy
~220
Current
farmers-market.org.uk / Pedddle
Easy
102 / 7 food stallholders
Current
NMTF / NCASS
No list
25k members (NMTF)
—
Nextdoor, findyourbaker
Easy by town
Unknown / 109
Variable
Trainers, insurers, wholesalers
A few named case studies; no lists
<10 named
—
[inference] Where the effort goes: a nightly FHRS diff (new FHRSIDs in 7841/7846/7839/4613 with the private-address pattern) is the only source that is national, fresh, free and lawful to hold; use it to size, to target the paid search and the launch social, and to feed post-to-shops and in-person market work — not electronic outreach. The market directories choose which markets to walk, not whom to write to.
A2. Which moments make the need acute
Bottom line. The two triggers that are both findable and timely are first council registration (104,000 registrations in 2025/26, roughly a third from home kitchens, every one told by GOV.UK to provide allergen information and expect an unannounced inspection [verified]) and the calendar (the Christmas run, Mother's Day and Easter, Allergy Awareness Week, and 1 October itself — the fifth anniversary of Natasha's Law commencing), because both are predictable and reachable through search and public pages rather than through any list of people. Enforcement stories, ingredient-level recalls and marketplace rule changes are sharper but arrive at random and identify no individual we may lawfully contact, so they are content moments to post on the day, not targeting signals.
Scoring: Findable = can we reach the business at that moment without a list, cold email or DM (1–5); Timely = how close to the moment of need a lawful touch can land (1–5).
A2.1 First council registration — Findable 4 / Timely 3
GOV.UK Starting a food business (updated 19 Aug 2026): register "at least 28 days before you start trading"; "an inspection will be arranged if required", higher-risk first, "some may wait several months", typically unannounced; you must "provide allergen information to your customers", for distance sales "before the purchase of the food is completed" and "when the food is delivered"; keep HACCP-based procedures and supplier traceability records for inspection [verified]. The page does not say PPDS or Natasha's Law in the parts returned [verified, absence]. FSA September 2026 Board, Local Authority Performance Update: ~590,000 establishments (April 2026); 104,000 new registrations in 2025/26, up 8.7%; 36,000 premises awaiting first inspection (6.1%), down from 77,000 in April 2021 but "5,000 premises above pre-Covid levels" [verified]. FSA press release 24 Feb 2022: 37% of new ventures since March 2020 operate from domestic kitchens [verified]. None of four council registration pages read (Oldham, Worcestershire, Brighton & Hove, North East Lincolnshire) mentions PPDS [verified, absence]; Brighton & Hove warns an incomplete SFBB diary "could affect your score" [verified]. What inspectors are told to check: the FSA's Feb 2026 audit of Rhondda Cynon Taf sets out the Code of Practice expectation to "verify incoming traceability and supplier allergen information" and assess labelling, and found the LA's allergen assessments "poor quality" with unsafe samples handled by "low level verbal advice" — "a serious failure" [verified]. So the expectation is that the inspector asks for supplier allergen information and labels; practice is uneven [inference]. Lawful reach: Google at the moment they type "register food business" / "what happens after registering"; council pages; content on "after you register: the four things the inspector will ask to see". FHRS AwaitingInspection in aggregate only.
A2.2 Seasonal peaks — Findable 4 / Timely 3
Dated compliance-calendar events: Allergy Awareness Week 20–26 Apr 2026 (Allergy UK) [verified]; the UK's first National Allergy Strategy launched in Parliament 20 Apr 2026 with recommendations aimed at catering and retail bodies [verified launch; content reported]; Anaphylaxis UK's bake-sale guidance (7 Aug 2025) tells home bakers to label every bake [verified]; 1 October is Natasha's Law's anniversary — the fifth on 1 Oct 2026 [verified]. Sales seasonality: British Baker reports a three-week Christmas period at nearly 10% of annual sales for one bakery and the Oct–Dec "golden quarter" as a "huge slice" [reported — 405]; Google Trends extracts put "birthday cake" at its 2025/26 high in Jan 2026 and "wedding cake" peaking Feb 2026 [reported, weak]. [inference] A hamper is the archetypal multi-item PPDS product (several home-made items in one pack, one label) — the Christmas-run message writes itself.
A2.3 Marketplace onboarding and policy changes — Findable 3 / Timely 3 (4 on the day)
TikTok Shop UKFood and Beverage Listing Guidelines (updated 25 Aug 2026): food is "Invite Only"; sellers must provide Food Business Operator registration (or FHRS rating) and a HACCP plan for manufacturers; "ingredients known to cause allergic reactions must be clearly indicated" on the product page and physical labels, with the 14 allergens highlighted [verified]. The June 2025 BBC investigation (sellers listing allergens as "not applicable" or ingredients as "spices") produced no announced TikTok policy change per CIEH (6 Jun and 3 Jul 2025) [verified]; whether the invite-only gate pre-dates it could not be established [reported]. The FSA's Future of Food Regulation (Sept 2026 Board) puts online platforms under "a dedicated discovery and scoping exercise" with recommendations in 2027 [verified]. eBay UK: register with the LA before selling food; labels must carry name, ingredients, allergens, storage, nutrition, place of manufacture [verified]. Etsy's policy pages: 403 (a search extract says accurate ingredient and allergen listings are required [reported]). Amazon UK: login shell only. Lawful reach: content on "what TikTok Shop's food onboarding asks for and how to produce the label it wants"; our own TikTok channel timed to policy updates; never DMs.
Named cases read on the council's own site [all verified]:
Date
Council / business
Facts
Outcome
4 Dec 2023
Oldham — Uncles Tea Hut
Packaged cakes sold with no labels, again at follow-up
£2,088 fines, surcharge, costs; "I hope this prosecution serves as a warning to food shops across the borough"
9 May 2025
Newcastle — Rasika Restaurant
"Mixed nut powder" on Peshwari naan was 100% peanut; two customers hospitalised; no PAL on the menu; 7 offences each
Company £800 + £320 + £1,000 costs; director and manager fined
30 Sep 2025
Pembrokeshire — CKs Supermarket
Undeclared sesame in multiseed bread, undeclared sulphite, found at routine inspection
£12,000 per offence, £36,000 + £2,000 + £2,849.95
22 Jan 2026
Barnsley — ATM Family Ltd
Labels not meeting requirements, allergen info not in English, registration out of date
£6,600
27 Feb 2026
West Sussex — Spice World, Crawley
Milk protein in "lamb" doner undeclared; meat misdescribed
Company £1,400; directors fined
Only Oldham is a PPDS-cake case our buyer would recognise as "me"; small-operator fines run hundreds to low thousands; the £36,000 case was a supermarket. The fact that lands is Rasika's: a customer hospitalised because a supplier ingredient was not what the label said — a supplier-declaration failure, which is ProvenBatch's exact ground [inference]. No national enforcement statistics for PPDS since Oct 2021 exist; the FSA's Sept 2026 update says England/NI food standards data is "currently unavailable" [verified]. Non-compliance rates that can be quoted: FSA Surveillance Sampling 2022-23 — 17 of 47 PPDS samples (36%) had undeclared allergens, a fifth of bread products, all PPDS [verified]; Retail Surveillance 2024/25 — bread 26% satisfactory, small FBOs 71% compliant vs 82% for large retailers [verified]; FSA PPDS evaluation (Apr 2023) — 24% not fully compliant, 17% moved foods out of PPDS to avoid labelling, 51% reported higher costs [reported — FSA pages now 410/404; via Roythornes, Oct 2023]. Inquests: Celia Marsh PFD report (7 Dec 2022) called for "a robust system to confirm the absence of the relevant allergen in all ingredients and during production" [verified]; Owen's Law remains voluntary guidance (5 Mar 2025), evaluation due spring 2026, law not before 2027–28 [reported]. Lawful reach: a same-week post via group admins in that borough and on our own channels — "what the Oldham case actually required" — reaches every business there while the story is live; no individual is targeted.
A2.5 The rest, ranked
Rank
Trigger
Findable
Timely
Volume
Lawful reach
Evidence
1
First council registration
4
3
~104k/yr; ~37% home
Exact-match search on registration queries; council pages; "after you register" content
Strong
2
Calendar (Christmas from late Oct; Mother's Day/Easter; Allergy Awareness Week; 1 Oct)
Ingredient-level recall (mustard/peanut Sept 2024 is the model: GOV.UK 30 Sep 2024 told every business to check; Anaphylaxis UK: "hundreds of products may be recalled" [verified]; ~84 allergy alerts in the year to Jun 2026 [reported, aggregator])
2
4
Rare at ingredient level
Same-day post: "if you bought X from Y, here is how to check your batches"
Moderate
7
Pending unannounced inspection
2
2
Every new business
As 5
Moderate
8
First wholesale / farm-shop request (legally moves the maker from PPDS to full prepacked labelling — Business Companion, June 2025 [verified]; retailer asks not verified)
2
2
Unknown
Farm shops' supplier pages; FRA
Weak
9
New FSA guidance (a labelling blog's "17 Jul 2026 new PAL rules" is contradicted by GOV.UK page dates — the technical guidance was last updated 23 Aug 2023; the 17 Jul 2026 change is a link update [verified])
3
3
Rare
Explainer on the day
Weak
10
"Is it nut free?" / school policy (GOV.UK school allergy guide 14 Sep 2026 [verified])
1
1
Constant, private
Message copy only
Weak
One forward signal: the FSA's Future of Food Regulation (Sept 2026) names "enhanced registration for all" as a core outcome, with Wales asking to explore "prior approval rather than a right of registration"; consultation June 2027, piloting from Jan 2028 [verified]. If registration becomes a gate, trigger 1 strengthens [inference].
A3. Who already has our customers' trust and could refer
Bottom line. The referrers with real trust and a working "recommend things" habit are (1) cake and food-business coaches and podcasts, who already run 30% affiliate schemes and take guest pitches, (2) insurers' knowledge centres, whose own FAQ wording ("labelled products accurately") is the buying trigger in the insurer's voice, and (3) shared kitchens and council food-enterprise centres, which already publish partner-perk lists and hand tenants to advisers. Councils and the FSA never link third-party tools, the national wholesalers have already chosen a caterer-facing labelling partner, and the trainers' affiliate programmes pay outward — so those three are audiences or data sources, not referrer rails.
Category
Realistic ask
What is in it for them
Named UK example with evidence
Grade
EHOs and council food-safety teams
Not a link — ten FSA/council/growth-hub/CIEH pages read, zero third-party links. Ask instead: a plain un-branded PPDS checklist they can hand out, tool named in the footer; and make the product EHO-legible so the inspection itself is the referral
Fewer failed allergen tests (Kent Scientific Services: ~11% of allergen tests failed, quoted on Kent & Medway Growth Hub, 10 Feb 2022)
Staffordshire Trading Standards reminded "home bakers, cake shed operators and market traders" about allergen labelling on 15 Sep 2026 — chasing our buyer this month, directing only to the council page
[verified]
Level 2 food-hygiene trainers
Two-way: join HST's/Virtual College's affiliate programme and link Level 2 from onboarding; then ask for a "what happens after your certificate: labelling" guest piece
Affiliate sales to them; content for their hub
High Speed Training on Paid On Results: 20% (1+ sales/mo), 25% (50+), 30% (100+), 60-day cookie, AOV £59, merchant since 14 Nov 2013; Virtual College via Awin 10%/5%, 30-day cookie; iHASCO referral/reseller partners
[verified]
Accountants for food micro-businesses
Get onto their "tools we recommend" list; extended trial for their clients
Cleaner client numbers; content they want
Dead Simple Accounting "Accountant for Caterers & Bakers" publishes a partner list (FreeAgent, Xero, Mettle, Tide, SumUp…); Livingstones Accountants sells a cake-shop package at £149.99/mo
[verified]
Packaging and label-stock suppliers
A "print-ready for your SKU" template and a link both ways; a guest compliance guide
Every label we print is stock they sell; SEO content
Magic Sparkles (edible glitter) published a 2,000-word Natasha's Law cake-labelling guide on 13 Sep 2026; Brother UK's PPDS page routes enquiries to "expert food labelling partners" and a Partner Programme (form is multi-outlet shaped); MUNBYN UK has affiliate/dealer programmes (inward); NIIMBOT UK has none; Good Cake Day sells a £77–129 Natasha's Law course with an affiliate scheme
[verified]
Commercial kitchens and enterprise centres
Perks listing (extended trial for members); one lunchtime "labels that survive an EHO visit" talk; line in the tenant onboarding pack
A tenant's labelling failure is the landlord's headline
Mission Kitchen Pro (£89/mo) lists partner perks — Raja packaging up to 50% off, Studio Mon 20%; Encore Kitchens (ex-Foodstars, 25 UK sites, 550+ kitchens, "300+ brand partners", Members Perks); Hartlepool Council's enterprise centre advises food tenants (4 Mar 2026); Broadland Food Innovation Centre (13 units, 100+ cluster businesses)
[verified]
Ingredient wholesalers
An Erudus integration so ProvenBatch reads the declarations wholesalers already publish — attacks the onboarding barrier directly; regional wholesalers without a named partner
NMTF member deal (their /deals list has a bank, the AA, solicitors, an insurer and an accountant — no software yet); NABMA "PPDS check for your food traders" via Market View; a line in operators' trader-acceptance emails
Operators carry the enforcement worry
NMTF /deals (Zempler, AA, Chafes Hague Lambert, FedInsure, APH, Harris + Co); NABMA takes corporate sponsors, conference 29–30 Sep 2026; The Makers Market T&Cs require "Allergen and ingredient information must be displayed" and current FHRS
[verified]
Insurers
A knowledge-centre guest article — "what 'labelled accurately' means, and how to prove it" — and a link exchange
Fewer disputed claims; SEO content
Protectivity home-bakery FAQ: product liability covers allergen incidents "provided you've taken reasonable precautions and labelled products accurately" (content dated 7 Jan 2026; from £4.43/mo; refer-a-friend £25/£25); Simply Business "How to start a home bakery" (updated 31 Mar 2026) names allergen labelling and no software; CMTIA £69/yr from 1 Jan 2026; PolicyBee from £11.93/mo
[verified]
Business-support bodies
BIPC "Become a provider": a free one-hour "labelling a home food business legally" session; refresh growth hubs' four-year-old Natasha's Law pages
Neutral stage; content
British Library BIPC network (London hub + six borough libraries; "Become a provider" link) [verified; national count reported]; Kent & Medway Growth Hub Natasha's Law page (10 Feb 2022); Enterprise Nation "Tools and offers" and "Work with us" (no price shown); FSB Marketplace [reported — JS page]
[verified unless marked]
Cake and food-business coaches
Podcast pitch form; co-branded update of Annie Bennett's free "Labelling Guide"; 30% recurring affiliate matching the CakeFlix norm
Free expert content; affiliate income; answers their members' #1 anxiety
CakeFlix affiliate 30% of initial membership/tutorials, 20% Master, 30-day cookie, £50 payout floor; Annie Bennett's links page offers a free Labelling Guide with no partner disclosed; Kate Tynan / The Cake Business Club "100+ cake makers", Coach of the Year 2025; The Business of Cake Making podcast has a public pitch-an-episode form; Good Cake Day affiliate via goaffpro
[verified]
Ranking for a solo founder (trust × reach × ease): 1 coaches and podcasts; 2 insurers' knowledge centres; 3 shared kitchens and enterprise centres; 4 NMTF/NABMA (no published partner process — needs a phone call); 5 BIPC. Councils, national wholesalers, trainers and printer brands score low as referrers because each would be asked to change policy rather than add a line [inference]. No UK insurer was found publishing "policy conditions require Natasha's Law compliance" as a formal condition; Protectivity's FAQ is the nearest [verified absence].
A4. What the closest analogues do
Bottom line. The self-serve UK analogues (FoodCore, AllergenKit, Stocksmith) grow on three near-zero-cost levers — free no-signup tools, a large programmatic content and landing-page estate, and a visible solo founder who answers personally — while the enterprise players (Nutritics, Kafoodle, Planglow) grow through wholesaler catalogues, demo-led sales and contract-caterer logos that we cannot and should not copy. Bake Diary's closure (email ~10 May 2025, last day 31 May 2025, no migration tool, domain now dead) is documented in two UK cake-business podcasts and several "alternative" pages, and the two things every replacement rushed to say — "your data is exportable" and "we are still here" — are the objection ProvenBatch must answer in its own copy.
Company
Pricing and trial (current)
How they acquire, as visible
Reviews and social
Grade
FoodCore (UK)
£25/£40/£65 inc VAT; 7-day no-card trial; "Nothing happens automatically" at the end
"FoodCore is me" — named solo founder, "I'm not a chef"; ~150 blog posts 2 Jan–19 Aug 2026; 5 vertical landing pages (home bakery, cake business, market stall, catering, kitchen); 4 "FoodCore vs X" pages; 4 free no-signup tools (recipe cost, margin drift, PPDS checklist, allergen matrix); Open Food Facts barcode import; Stripe/Shopify/WooCommerce; three homepage testimonials; no affiliate/partner programme
Free no-signup matrix builder, chart, 14-allergen poster, draft label maker; Natasha's Law guide with email signup; "Tell Finn one recurring bake and he'll reply personally"; UK-region database; "Cake sheds · home bakeries · market bakers"
No counts, no testimonials, no social links, no affiliate
[verified]
CompliChef (complichef.co.uk; the .com is a Peruvian caterer)
£39/£160/£170 per month on the homepage vs "£29 per site" on the switch page; Lite £79.99/yr; Solo Labels = £310 Sunmi printer + £30/mo; card taken via Stripe, 14 days
Chef-founder "Nick"; sector pages; a "Switch from Trail/Navitas/Food Alert" page; hardware partners; live counters ("728 food labels printed" as of 1 Jan 2026)
From $20/mo; 14-day no-card trial; "Your data is never deleted… export everything"
"10,000+ product businesses"; free costing spreadsheet; "Recipe Costing Software for Makers" comparison; /compare/bake-diary; affiliate 20% for 12 months, 90-day window, brand-bidding ban, aimed at bookkeepers, educators, newsletters
Capterra 4.6
[verified]
Bake Diary (IE; closed 31 May 2025)
€6.95/mo [reported]
Closure email ~10 May 2025 (Sugar Cookie Marketing FB post, 10 May 2025: "shutting down in 2 weeks"); comments: "Two weeks notice is waaaay short. I've been using it as a base for my taxes returns"; "2 weeks to transfer everything is almost impossible"; an 8-year user in wedding season; several never got the email [verified]. No migration tool. Covered by Slice of Success (13 May 2025, "already causing a lot of panic"; recommends Baking It, UK) and The Business of Cake Making Ep 162 (26 May 2025) [verified via RSS]. Domain dead/resold; app gone from the GB App Store [verified]. Founder name and stated reason: not found
Reviews not retrievable
[verified unless marked]
Others seen
Paddl £69/location, 30-day card trial, 12 free tools; BakeProfit (US) free tier + $9.99, "5,000+ bakers"; Castiron (US) raised $6M, shut late 2025 [reported]; CakeBoss $149 first year; BakerInbox (UK) £19/mo, 14-day no-card, published the UK "Bake Diary alternative" page 15 Jun 2026; Bake Boost (CA) extended its trial to 30 days for Bake Diary refugees; Allergen Checker (UK) Free/£20/£47; Erudus caterer plans £10/£17/£20/mo
[verified]
Cross-cutting patterns [verified unless marked]. Every self-serve product with traction runs no-signup tools, and the allergen-matrix builder appears three times (FoodCore, AllergenKit, Paddl) — crowded; a free PPDS label checker ("paste or photograph your label, we tell you what's missing") is offered by none [verified absence; inference that it is open]. Review counts are tiny everywhere (0–25); volume lives in app stores for consumer-style apps. Only Stocksmith publishes an affiliate programme. Paid-ad activity could not be verified for any analogue (Google Ads Transparency is a JS shell; Meta Ad Library 403). No founder interview surfaced for FoodCore, AllergenKit, Bakesy or Bake Diary.
Copyable at near-zero cost [inference on verified facts]: (1) a free no-signup PPDS label checker; (2) Stocksmith's end-of-trial wording plus a "what happens if ProvenBatch closes" page; (3) a UK "Bake Diary alternative" page stating pounds, UK-held data, Natasha's Law labels; (4) concierge first-product onboarding as a public offer (AllergenKit's "tell Finn" made ours); (5) five to seven vertical landing pages including farm shop and butcher, which FoodCore lacks; (6) a factual "vs FoodCore" page — nobody writes one; (7) Erudus purchase-history import for caterers, delis and farm shops; (8) podcast guesting — two UK cake podcasts devoted episodes to a software closure, so software is on-topic; (9) affiliate terms modelled on Stocksmith; (10) directory listings for the badges, expecting no reviews.
Off the table under §4, which the analogues do: bakery-only branding (AllergenKit, Bakesy, BakerInbox); LinkedIn and X (Nutritics, Kafoodle, CompliChef, Planglow); own Facebook groups and group posting (Bake Boost); testimonials before real ones exist (FoodCore, BakeProfit); demo-led funnels with account managers (Planglow, Nutritics); hardware and label lock-in (CompliChef, Planglow).
B. Reaching them for conversion
B5. Message — testing "compliance fear beats time saved beats cost control"
Bottom line. The hypothesis is half right: fear is the trigger for a first-timer and for anyone who has never labelled, but for a business already labelling by hand the evidence (the FSA's own 2019 consultation, competitor users' review language, the 2023 evaluation) says the live pain is keeping labels correct when recipes and supplier packs change — an error-and-time problem, not a fear problem — and cost is a barrier to software rather than a motive to buy it. The angle the hypothesis misses is customer trust: allergic consumers say in the FSS's own research that they avoid small outlets because they do not expect or trust a full ingredients label, and the FSA's sampling shows that when small businesses fail it is mostly because there is no label or no ingredients list at all, not because the list is wrong.
B5.1 What businesses and consumers actually say (verbatim where a primary page allowed it)
Verbatim Facebook, Mumsnet, TikTok and YouTube comments could not be captured this pass (those sites returned 403/404/CAPTCHA and the search budget ran out); the closest primary substitutes are below and the gap is listed at the end.
Who
What they said
Source
Grade
Businesses responding to the FSA's 2019 PPDS consultation (126 business responses, 67% micro/SME)
Concerns listed: "Cost and financial burden, particularly for small/micro businesses"; "Time and resources needed to source, print, and update labels"; "Risk of mislabelling in busy kitchen environments"; "Complexity of ingredient substitution management". Only 13% of businesses preferred full-ingredient labelling; 41% preferred "ask the staff" stickers; 73% of individuals supported full labelling
91% aware; 68% "have all the information needed"; ~50% report higher costs; "set-up costs were significant; ongoing costs… did not pose an issue to the survival of the business"; more precautionary "may contain" as an unintended consequence
FSA blog 19 Jul 2023 (report chapters now 410)
[verified blog]
Same, via trade press
24% "not fully compliant"; 72% started applying PAL; 17% started selling previously PPDS foods as non-prepacked; 75% of LAs want more action/training
Food Manufacture 2 Oct 2023; Roythornes 2023
[verified secondary]
Customers' questions to Planglow, Sep–Oct 2021
"What about micro-businesses, do they need to comply?"; "Is there an easy way to test whether it is PPDS?"; "What would happen to labelling if the person managing it was off work?"; "Will EHO be checking Natasha's Law during visits?"; "What about plated meals & made to order?"; cakes "bagged up when the customers choose"
planglow.com Natasha's Law FAQs
[verified]
Our own beta-request form (9 completed rows, Aug–Sep 2026)
Labels today: Canva 5 of 9, Word/Pages 1, "supplier or printer" 2, none 1. Pains typed: "copy and pasting all my ingredients to form a label"; "Having to input every ingredient, it's very time consuming"; "manually typing them out before printing"; a chocolatier: brands change "frequently due to local stock issues"
production beta_requests
[repo]
Our own in-app feedback, Sep 2026
"this doesnt make any sense to me??"; "i created a recipe but it wasnt reading the ingredients properly so i deleted it… now it wont let me add it?"; the receipt scanner "prompts to take photos of declarations you don't have… it doesn't tell you which ones!"
production feedback_items
[repo]
Kafoodle users (Capterra, 18 reviews)
"It's a long process to set up when filling ingredients yourself"; recipe changes "doesn't always flow through to the label tool immediately"
capterra.com
[verified]
Nutritics users (Capterra, 25 reviews)
"Data Entry is painful, clumsy and slow" (2021); "Buggy, horrendously slow" (2 Jul 2025)
capterra.com
[verified]
Bakesy UK users (App Store, 87 ratings)
"it asked me to pick a subscription right away before even letting me see the app features" (10 Sep 2025); "I've found it easier and quicker to do it the old fashion way"
Stall: "I was terrified of getting Natasha's Law wrong"; bakery: "an hour updating our labels whenever we changed a recipe"; meal prep: "35 products… Keeping allergen information accurate manually was impossible"
foodcore.io
[verified page; claims unverified]
Allergic consumers, Scotland (FSS qualitative research, 44 participants, Mar 2023)
"The lack of a full ingredients label means I can't risk it"; "I would like to buy more locally, but I don't trust them because you cannot rely on the content of the food matching the labelling… they change their ingredients often and they don't keep up with the labelling"; "It just gives me more confidence and trust in that retailer [if PPDS foods are labelled]"; "I just avoid anything that has 'may contain' warnings"; "Shops are just covering themselves by using 'may contain'"
FSS PPDS consumer research PDF, June 2023
[verified]
Councils, this month
Staffordshire Trading Standards, 15 Sep 2026: "home bakers, cake sheds and market traders are being reminded to check allergen labelling requirements"; requirements "apply to all business sizes, including home-based operations"; 11 May 2026: "will take firm action against those who fall short"
alittlebitofstone.com
[verified]
A supplier, this month
Magic Sparkles (edible glitter), 13 Sep 2026: "A shimmer that took thirty seconds to apply can be the exact thing missing from an otherwise correct label"; "Overlabelling carries no penalty. Underlabelling does."
magicsparkles.com
[verified]
The numbers that can be quoted [all verified]. FSA Our Food 2023 (8 Oct 2024): 47 PPDS foods sampled, 17 (36%) had an undeclared allergen — "either there was no label present or the allergen was not listed"; "Compliance failures in the sampled products were restricted to smaller food businesses. A high percentage were due to the absence of labelling, or some labelling being present without an ingredients list rather than the ingredients list being incorrect." Allergy alerts: 77 (2020), 83, 83, 64 (2023), 101 (2024); milk the most common undeclared allergen most years; 2024's peanut-in-mustard incident produced 34 alerts across 59 brands (Our Food 2024). Food and You 2 Wave 11 (31 Mar 2026): 23% of adults report a food hypersensitivity; 59% had a bad reaction (up from 42% in 2021). Confidence in allergen information by channel (Wave 2, 2021): cafés/sandwich shops 79%, takeaway direct 63%, delivery apps 50%, Facebook Marketplace 21%. FSA out-of-home research (7 Oct 2025): 87% trust written information vs 75% verbal; 25% "couldn't see allergen information" at their last purchase. FSA SME research (3 Oct 2024): updating written allergen information "may carry costs… meaning they are unable and/or unwilling to do this as regularly as needed". 2019 Impact Assessment: a typical outlet spends "approximately £100.00 annually on labels".
The law, stated so it can be quoted safely [verified from legislation.gov.uk and GOV.UK]. PPDS = "food that is packaged at the same place it is offered or sold to consumers and is in this packaging before it is ordered or selected", including "some food sold at mobile or temporary outlets" (GOV.UK, updated 17 Jul 2026). Failing to comply with reg 10 of the Food Information Regulations 2014 "constitutes a criminal offence" with "a potentially unlimited fine" (FSA technical guidance; level 5 uncapped in England and Wales by LASPO 2012 s85 since 12 Mar 2015; Scotland/NI scales not checked). "May contain" only after a risk assessment finds "an unavoidable risk… that cannot be sufficiently controlled" and "could be considered misleading if it doesn't reflect genuine risk". Labels "can be handwritten" if legible (GOV.UK caterers guide, 17 Jul 2026) — so "hand-written is illegal" is a claim we must never make. No prosecution of a home baker or market stall for PPDS labelling was found; do not imply one.
B5.2 Which framing converts whom [inference on the verified evidence]
Nervous first-timer (home baker, new stall, first market season). They ask "does it apply to me?", "are cupcakes in a box PPDS?", "do allergens have to be in bold?", "will the EHO check?". Framing: acknowledge the fear, then remove the guesswork with facts. Not "you could be fined" (they know; it paralyses) but "here is the rule, here is the label, here is the proof":
"A cake in a box on a stall is PPDS. The label needs the name and the full ingredients with the 14 allergens emphasised. ProvenBatch builds it from the packs you actually used."
"Most labelling failures the FSA found in small businesses were missing labels or missing ingredients lists, not wrong lists." (Doing something correct beats doing nothing.)
"'May contain everything' is not a label; the FSA says it must follow a risk assessment."
The trial terms as the fear-reducer: 30 days, everything unlocked, no card, read-only at the end, nothing deleted.
The onboarding promise, because that is where they die: "Photograph the pack; the declaration is read for you." Kafoodle's "long process to set up when filling ingredients yourself" is the objection to pre-empt.
Established business already labelling by hand (Word/Canva/Avery, a Brother or Dymo). They are not afraid; they are tired, and their risk is a stale label. Framing: "The label follows the batch." Change a supplier pack, swap flour brands, add glitter: the label regenerates and the record shows which pack went in. Secondary: "records your EHO can read" and "recall watch against the packs you actually use" (the 2024 mustard incident sat in bought-in ingredients, not recipes). Cost last: "£9 a month, about what the FSA reckons a small outlet spends on label stock in a year." Do not lead with "criminal offence" for this group; the FSS consumer quote about small outlets that "change their ingredients often and don't keep up" is the better mirror.
Revised order: first-timers: fear → resolved by "know it's right"; established: change-proofing (time + error) → customer trust → cost.
B5.3 Message angles per vertical (plain voice; the fact behind each is verified unless marked; the angle's effect is inference)
Home bakers / cottage bakeries. (1) "A boxed cake, cupcakes in a box, brownies bagged for a stall: all PPDS" (GOV.UK bakers guide 2021; FSS example). (2) "Working from home does not exempt you. Councils are writing to home bakers and cake sheds by name" (Staffordshire TS, Sep 2026). (3) "The decoration is part of the label" (Magic Sparkles; milk the most-recalled allergen). (4) "Stop writing 'may contain all 14'. It is not a label, and allergic customers walk away from it" (FSA technical guidance; FSS quotes). (5) "If you used Bake Diary, this is the sub-£10 slot with Natasha's Law labels built in" (Bake Diary €6.95 [reported], closed 31 May 2025 [verified]).
Caterers / meal-prep. (1) "Packed lunches, meal boxes, soup in pots and enclosed platters are PPDS; open platters and made-to-order are not" (GOV.UK event caterers guide, 17 Jul 2026). (2) "Thirty-five meals a week with different ingredient lists is exactly the workload manual labels fail at" (2019 consultation: "complexity of ingredient substitution"). (3) "One supplier swap should not mean re-typing every label" (Kafoodle review). (4) "Your customers trust written answers more — 87% vs 75%" (FSA Oct 2025).
Butchers. (1) "Sausages, burgers, marinated steaks and stir-fry packs packaged before the customer asks are PPDS; loose meat in the counter is not" (GOV.UK butchers guide, 17 Jul 2026). (2) "A butcher's label also needs the meat percentage" (same guide; confirm QUID support before using [inference]). (3) "A regular customer asked her butcher to 'surprise' her with a sausage and had a reaction" (FSS 2023) — a story, not a threat. (4) "Sulphites and rusk are the quiet ones" [inference from practice, not a sourced stat].
Delis / farm shops. (1) "Cheese, olives, chutneys and cooked meats you portion and wrap before sale are PPDS; sliced to order is not" (GOV.UK introduction and butchers guide; no deli-specific guide exists [verified absence]). (2) "Allergic customers want to buy local and mostly don't" (FSS: "I would like to buy more locally, but I don't trust them"). (3) "Trading Standards may be the inspector here" (Business Companion, Mar 2025: a folded-over bag is prepacked, an open bag is not). (4) "Bought-in prepacked is the maker's label; made-and-packed here is yours" (FSA technical guidance ¶95).
Cafés / sandwich bars. (1) "Prepacked sandwiches, salad boxes, pots of soup and packaged pies in the chiller are PPDS; the open salad bar and unpackaged cakes are not" (GOV.UK restaurants/cafés/pubs guide, 17 Jul 2026). (2) "Natasha's Law came from a baguette" (NARF). (3) "Cafés are already trusted (79%); a missing label is what breaks it" (Food and You 2). (4) "The FSA's stated worry is small businesses that don't keep labels up to date when the kitchen changes something" (FSA SME research).
Market stalls / street food. (1) "If you pack it at home and sell it from the stall, it is PPDS. 'Made to order' only covers what you put in a box after they ask" (GOV.UK mobile sellers guide; NCASS). (2) "Trading Standards reminders now name market traders" (Staffordshire). (3) "Buyers trust a marketplace seller far less than a café (21% vs 79%); a proper label closes that gap" (stall ≈ marketplace is inference). (4) "Labels can be handwritten if legible; the problem is forty lines at 6am."
Chocolatiers. (1) "Milk is the most frequently undeclared allergen in UK recalls" (Our Food 2023/2024). (2) "Hand-packed boxes and bars for a stall, fair or counter are PPDS; a pick-and-mix filled to order is not." (3) "'May contain' must follow a risk assessment, and coeliac and allergic buyers avoid it." (4) "The names inside your couverture — soya lecithin, milk powder, nuts — belong on your label, in the right order" (FIC Art 9(1)(c) via FIR reg 10).
Preserves / sauces. (1) "Jars filled and sealed before sale are PPDS wherever you sell them; the rule follows the jar." (2) "Sulphites, celery and mustard in chutneys are the quiet ones; mustard was at the centre of 2024's biggest allergen recall" (Our Food 2024). (3) "A recall against an ingredient you bought is your problem too" (FSA incidents 2024/25: labelling and cross-contamination the main causes). (4) "Jars you supply to a farm shop are prepacked, not PPDS — they need the full label, which is the same batch record" [inference on FSA scope].
Phrases the evidence supports: "criminal offence"; "potentially unlimited fine"; "applies to all business sizes, including home-based"; "packaged before it is ordered or selected"; "may contain must follow a risk assessment"; "most small-business failures were no label or no ingredients list"; "87% trust written information". Avoid: "hand-written labels are illegal"; "£5,000 fine" (out of date in E&W since 2015); any home-baker prosecution story (none found); vendor testimonials as proof; "Natasha's Law says…" for anything beyond FIR/FIC.
B6. Channel sequencing for a solo founder — 90 days from 1 October
Bottom line. With roughly ten hours a week and £150 a month, the order is: fix activation first (because the benchmark says 4–8% of opt-in trials pay and the beta says onboarding is where they die), then coach and admin partnerships (the only proven source of beta signups), then exact-match search plus the four pages that catch people at the registration moment, then the Metricool-run social calendar with a reactive trigger routine, then the institutional rails and the referral credit; everything is gated at week four on one measured leading indicator and killed or cut to maintenance if it misses. Nothing here needs daily manual posting or any spend beyond the £150 envelope [inference throughout, built on §A and §B evidence].
Precondition (week 0, ~3 hours, £0): carry utm_*/gclid from the site into the signup record, add the "how did you hear about us?" field, and issue one promo code per partner — without these no channel below can be judged (the 15 Sep Metricool research already says so [repo]).
Order
Channel work
Hours/wk
£/month
Leading indicator (read every Friday)
Week-4 gate
1
Trial activation + concierge first product (§B8): one-product first session, behaviour-triggered service emails, "we'll do your first product with you" for anyone with no pack photographed by day 1
3
£0
% of started trials with a real label from a photographed pack within 7 days (target ≥30%); concierge offers accepted and their conversion
If <15% activated, pause channels 3–4 and put their hours here; trials into a leaking funnel are wasted
Sessions agreed; trials attributed by promo code / "how did you hear" (target ≥5 per session)
If no session agreed by week 4, drop to one follow-up a month and move the hours to channel 3
3
Search: exact-match Google + four pages — "Bake Diary alternative UK", "ProvenBatch vs FoodCore", the free PPDS label checker, "After you register: what the inspector will ask to see"
2
£60–100
Exact-match impressions/week (if <300, the volume is not there); cost per started trial at £100 spent (≤£20 keep; >£40 kill) — and note §0.7: at 4–8% conversion even £20 is £250–500 a customer, so the honest target is ≤£8
Kill paid at >£40/trial after £100; keep the pages (they compound)
4
Owned social via Metricool (Starter): the queue plus a reactive routine — a post within a day of a council enforcement release, an ingredient recall or a platform policy update (§A2); one £25 boost of the best launch Reel
1.5
£17 + £25
Link clicks to /trial per week (target ≥20 by week 4) and ≥1 attributed trial/week
If missed, cut to Autolist recycling and the reactive posts only
5
Institutional rails (§C12 leads 3, 4, 5, 7, 8, 10, 15, 20): one guest article, one perks listing, one member deal, one council page, the Erudus partner conversation
1
£0
Replies received; one live listing or article by day 60
If nothing live by day 60, keep only Erudus and Protectivity warm
6
Referral credit (§B9): build "give a month, get a month" once (weeks 3–4, ~4 h), show it on the read-only banner
0.5 avg
£0
Referral codes used per 100 active accounts; read-only reactivations
None — it is a one-off build; report monthly
7
Cake International visit, 6–8 Nov (§C12 lead 18): meet OLBAA (stand C40), the Sugarcraft Guild, coaches; trial cards only where a stand invites it
1 day
£30–50 travel
Conversations that turn into a §C12 follow-up
n/a
Total: ~10 hours a week; £100–150 a month. Off the table and not in this table: daily manual community posting, cold email or DMs, exhibiting, any purchased list, any channel outside Facebook/Instagram/TikTok/YouTube.
B7. Lawful outreach mechanics — sole trader vs limited company
Bottom line. For a sole trader (or an ordinary England/Wales/NI partnership) the only first-contact routes that need no prior consent are a live phone call to a number not on the TPS, an addressed letter, and a face-to-face conversation; every email, SMS, WhatsApp or social DM needs consent or a genuine soft opt-in, and the ICO's own example says a free-trial signup is "negotiations for a sale" [verified]. For a limited company, LLP or Scottish partnership the PECR consent rule does not apply to email/SMS/DMs, but you must identify yourself, give a working opt-out, honour it, and satisfy UK GDPR (legitimate interests with an assessment) for any named person's address [verified]. The ICO's direct-marketing guidance is current as of 28 April 2026; the Guide to PECR pages are still "under review" and the small-organisations update is pending.
B7.1 Which guidance is current
Document
Status and date
URL
ICO Direct marketing guidance (Identify → Plan → Collect → Respect; Annexes A/B)
Published 5 Dec 2022; updated 28 Apr 2026 "to reflect the six month commencement schedule of the Data (Use and Access) Act" [verified]
DUAA commencement (legislation.gov.uk) [verified]: 19 Jun 2025 Royal Assent; 20 Aug 2025 (SI 2025/904) the definition of "direct marketing" inserted into PECR reg 2 — "the communication (by whatever means) of advertising or marketing material which is directed to particular individuals"; 5 Feb 2026 (SI 2026/82) charity soft opt-in (reg 22(3A)), PECR enforcement moved to the DPA 2018 penalty regime (s.115 + Sch 13) — ICO statement 5 Feb 2026: fines "up to £17.5 million or 4% of global turnover under PECR"; recognised legitimate interests (Art 6(1)(ea)
interests do not include direct marketing** (Annex 1 lists only public-task disclosure, security, emergencies, crime, safeguarding) [verified]; direct marketing sits in Art 6(11) as something that "may" be a legitimate interest — the ICO: "You must still do the three-part test" and "If PECR require consent, you must not use legitimate interests" [verified].
B7.2 Who is an individual subscriber (this decides everything)
PECR reg 2 [verified]: individual = "a living individual and includes an unincorporated body of such individuals" — sole traders and ordinary (E/W/NI) partnerships; corporate subscriber = companies, royal-charter bodies, Scottish partnerships, corporations sole, "any other legal entity distinct from its members" — Ltd, LLP, public bodies. ICO (28 Apr 2026): "individual subscribers (people, sole traders, ordinary partnerships)" vs "corporate subscribers… limited companies, LLPs and Scottish partnerships" [verified]. Scotland: a sole trader is an individual anywhere; only the Scottish partnership differs [verified]. Generic addresses: "The rules… apply even if you only hold a generic or role-based address" — info@ at a sole trader is still an individual's [verified]. "If this is not clear, assume they are an individual." (Guide to PECR, Using marketing lists) [verified]. Practical rule [inference]: without a Companies House entry or "Ltd/LLP" on their site, treat every home baker, stall, jam maker and caterer as an individual — which is the brief's own rule.
B7.3 Route by route
Route
Sole trader / ordinary partnership
Ltd / LLP / Scottish partnership / public body
Common to both
Live call (reg 21)
No consent needed unless the number is on the TPS or they previously objected [verified]
Same, against the CTPS [verified]
Screen both registers within 28 days of calling (registrations take 28 days to bite [verified]); present a CLI you can be contacted on (reg 21(A1)); "say who is calling… provide contact details or a Freephone number if asked" [verified]; legitimate interests + assessment; keep a do-not-call list. TPS/CTPS licence costs: not verified (site did not render)
Post
Lawful — "direct marketing by post is not covered by PECR" [verified]; UK GDPR legitimate interests with the three-part assessment; a named sole trader's trading address is personal data [inference]
Lawful; address a role or use LI for a named person
Say who you are, why you have their details and the source; the right to object is absolute and free ("There are no reasons that you can use to refuse") and must be told "at the latest at the time of your first communication" [verified]; check the MPS ("you should check, although it is not a statutory one") [verified]; suppression list
In person (markets, shops, trade counters)
Lawful — a conversation is neither electronic mail nor a call [inference on reg 2]
Lawful
UK GDPR applies once you write a name down; what you later send electronically needs consent captured on the form: separate unticked box per channel, named sender, purpose, how to withdraw, privacy-notice link (ICO's good example: "☐ I would like to receive marketing from you about your services by text message."; bad: the same without the channel) [verified]
Existing-customer soft opt-in (reg 22(3))
Applies: details obtained "in the course of the sale or negotiations for the sale"; "similar products and services only"; a free opt-out "at the time that the details were initially collected" and in every message [verified]. ICO: "A person doesn't need to actually buy anything… It's enough if 'negotiations for the sale' took place" — examples: "signing up to a free trial of your product or service", "requesting a quote", "asking for more details"; but it "must be an express communication… and it must involve them buying your products or services" [verified]
Not needed
A lead-magnet download is not addressed by the ICO and probably does not qualify (the competition example and "must involve them buying" point against) — put a separate unticked consent box on every download form [inference on verified text]. "There is no such thing as a third-party marketing list that is compliant with the soft opt-in" [verified]
Email / SMS / WhatsApp / social DM (unsolicited)
Consent or soft opt-in only — "electronic mail" covers "email and text (SMS) messages; picture or video messages; voicemail; in-app messages; and direct messaging on social media" [verified]; consent must be per channel ("simply saying 'electronic mail' is not specific or informed enough") [verified]
Permitted: "You can send unsolicited electronic mail marketing to corporate subscribers without consent or a soft opt-in" [verified]
Reg 23: never disguise identity; give "a valid address to which the recipient… may send a request that such communications cease" [verified]; for a named work address UK GDPR applies — legitimate interests + assessment, absolute right to object, tell them the source "at the latest within a month" [verified]; keep a do-not-email list
Reply where they asked in public
A public reply to a public question is not electronic mail and is not "directed to particular individuals" unless you tag or DM them to advertise [verified definition; inference on application]; "Electronic mail marketing is solicited when someone specifically asks you to send a particular message or type of information" [verified]
Same
A DM to someone who asked a general question in public but did not ask you is unsolicited → needs consent/soft opt-in for an individual. Public profiles are not a free source: "because someone's social media page has not been made private… doesn't mean that you are free to use their personal information for direct marketing" [verified]
DM to a business that messaged first
Your reply with the information asked for is solicited — fine; a later unrequested follow-up is unsolicited [verified]; if their message was "negotiations" (asking about price, plans, a trial), the soft opt-in covers follow-ups only if you offered an opt-out then and in every message [verified]
Permitted with identity + opt-out
Practical: end the first reply with "Happy to keep you posted on ProvenBatch by message — say 'no thanks' any time and I'll stop", and log it [inference]
B7.4 Consent wording and the record the ICO expects [verified unless marked]
Wording: an affirmative act ("Silence, pre-ticked boxes or inactivity should not constitute consent"); "the name of your organisation"; "why you want the data"; "what you will do with the data"; "that people can withdraw their consent at any time" and how; "prominent, concise, separate from other terms and conditions, and in plain language"; "a separate opt-in for each" purpose and channel; confirmation of having read a privacy policy is not consent; consent "is specific to receiving electronic mail marketing to a particular number or address". Model wording assembled from those requirements [inference; not an ICO template]:
☐ Yes, email me from ProvenBatch (David Biley trading as ProvenBatch) about PPDS labelling tips, product updates and offers. ☐ Yes, text me the same. You can stop at any time using the unsubscribe link in every email, by replying STOP, or by emailing hello@provenbatch.co.uk. Privacy notice: provenbatch.co.uk/privacy.
Records (ICO "How should we obtain, record and manage consent?" plus the direct-marketing guidance): who (name or identifier); when (dated document or timestamp; for oral consent a note made at the time); how (a copy of the capture form; for oral, the words used); what they were told ("a master copy of the document or data capture form containing the consent statement in use at that time"); withdrawal and when; plus "whether the customer is an individual or a company"; which method applies — consent or soft opt-in ("a simple set of flags or preference fields in your system is usually enough"); the source; channels consented to; objections and opt-outs; a suppression list checked before every send. Refresh: "consider refreshing consent every two years"; third-party consent older than six months is not usable. Privacy information at the point of collection: purposes, sharing, the right to object, and the UK GDPR Art 13 list (identity, lawful basis, recipients, retention, rights, ICO complaint route).
B8. Trial to paid
Bottom line. The best current primary benchmark (ChartMogul with Kyle Poyar, Jan 2026, 200 self-serve B2B products; US/global, so discount it) puts an opt-in, no-card trial at 4–6% trial-to-paid as "good" and 10–15% as "great", median 8% across models; card-required trials run ~5× higher but from fewer signups, so revenue per visitor is similar [verified]. Trial length barely moves conversion in the only randomised experiments that exist, whereas reaching a real first-value event in the first week does, so ProvenBatch's lever is not 14-vs-30 days or a card wall but getting one real PPDS label out of one real product inside the first session, with a done-for-you fallback for anyone who has not photographed a pack by day 2 [inference on verified evidence]. No UK trial-conversion figure exists anywhere found (FreeAgent, Crunch, Xero publish none).
The circulating "25% / 60%" averages have no traceable primary source (Lincoln Murphy's 2014 post does not contain them) [verified absence]; the "72-hour" activation statistic is unsourced in every page that repeats it — the nearest real evidence is ChartMogul's finding that "conversions peak in week 1" (2,500 companies, 2025) [verified]. Do not quote 25/60 or 72 hours. Trial length: 14 days is 62% of products, 30 days 14% [verified]; the two randomised experiments found a 7-day trial beats 14/30 by +5.59% relative (Yoganarasimhan et al., large SaaS firm) and a 7-day vs 3-day trial had no immediate effect but +20.9% total conversion over two years for a "task-oriented" product (Zhang & Duan 2025, 680,588 users) [verified]. Keep 30 days; treat days 0–3 as the trial that matters [inference].
Activation. Median activation rate 25% (Lenny's 2022, 500+ products); "healthy" 25–35% (Appcues, Jun 2026) [verified]. Method: list 3–5 candidate first-value events, correlate each with 30/60/90-day retention, set a window (most use 7 days), validate by interview [verified]. Checklist completers converting ~3× (Sked Social) and HubSpot's five-actions-in-two-weeks 3× are reported only. Concierge/done-for-you: ConvertKit's free concierge migration is credited with $1.3k → $5k MRR in six months and ~1.5% vs ~5.5% churn for migrated accounts [reported — Failory; Kit's own page 403]; Podia moves "products, courses, and people by hand" free on every plan [verified]; Groove's founder welcome email drew a 41% reply rate [reported]. No controlled measurement of a setup offer at a sub-£40 price exists [verified absence]. Read-only vs lock-out: Basecamp freezes ("data will remain intact, though inaccessible"), Todoist and Notion downgrade to free, HEY recycles the address after 90 days [verified]; no published data compares read-only against lock-out; a read-only end state is functionally a reverse trial, so ChartMogul's 4–6% / 8–12% is the best prior [inference]. Under PECR, service messages are outside the rules but "if your service message has elements that are direct marketing… it will count as direct marketing" [verified].
B8.2 Applied to ProvenBatch's barrier — proposals [inference, built on the shipped product as described in outputs/USER-GUIDE.md]
Activation definition: a real label produced from a batch of a real product whose recipe uses at least one supplier product with a declaration read from a pack photo, within 7 days of signup ("real" = not the demo data). Instrument the chain the Today card already draws: pack photographed and declaration accepted → recipe with quantities made a Product → batch with packs confirmed → label printed. Report weekly the share reaching each step by day 1/3/7/30. Target ≥30% activated by day 7 and ≥50% of activated trials paying — which at 30% activation is ~15%, the "great" band. Re-run the correlation once ~100 trials have aged past day 30.
First session — "first label in ten minutes": possible only if the session is scoped to one product and only the packs it needs (typically 4–8). Lead with "Which product will you label first?", take the recipe (typed, pasted or photographed), generate the pack list from the recipe lines and walk the camera through them ("Pack 1 of 6: photograph the flour bag's ingredients panel"). That turns "enter dozens of declarations" into "photograph six packs". Keep the demo off the critical path. Show the label taking shape as each declaration lands (a preview of derivation, not a printable label — it must obey the labelling design document). Make the drop-pile the second door for people who already have pack photos on their phone.
Lifecycle emails under PECR: at signup an unticked "Send me occasional tips and product news" box (consent; do not lean on the soft opt-in even though it is available — an untick-to- refuse pattern is what the brand said it would not do). Service-only sequence for everyone, each message about the account and free of upgrade pitches: day 0 "Your first label: which product?"; day 1 only if no pack photographed — the concierge offer; day 3 if pack but no label — the exact next step; day 7 on first label — "here's what the EHO can now see", or the concierge offer once more; day 23 and 29 — the existing 7-day and 1-day notices, stating plainly what read-only means; day 30 — the existing read-only notice; day 37 and 60 — one post-expiry note each, then silence. Plan comparisons, annual discounts and feature news go only to the consented cohort. Behaviour triggers over dates; a Supabase scheduled function on the milestone table is enough.
Concierge wording: "We'll do your first product with you — free, during your trial. Photograph the packs for one product you sell (most need five to ten) and drop them on Today, or email them to your account's receipts address. We'll read them, check every declaration against the pack, and set up the recipe so your first label is waiting for you to approve. Nothing is printed until you've checked it. Usually done the next working day." Offer at day 1 and day 7 to non-activated trials only; cap per week; measure accepted vs declined; automate or stop once the sample says so. Rob Walling's "refund month one if you complete onboarding" is the reserve.
Read-only end state: "Your trial has ended. Nothing has been deleted and nothing has been charged. Everything you recorded is still here: you can open any batch, view and export any label or record, and show an inspector exactly what went into anything you made. What's paused is new work: recording a batch, printing a label for it, adding a pack. Choose a plan (from £9 a month) and it all resumes where you left off." Past labels and batch records stay viewable and exportable (they are the EHO evidence); reprinting for a new batch is new work and stays paused. Report read-only reactivations at 30/60/90 days as their own funnel line.
B9. Referral and word of mouth
Bottom line. Ship a two-sided "give a month, get a month" credit first — no cash, no affiliate software, no tracking beyond a code in Settings and a Stripe credit — because every cash or commission scheme read (Xero, Tide, Octopus, Stocksmith, Metricool) needs a qualifying period, fraud controls or $39–99/month of tooling, and the only one that fits a solo founder at £150/month is the FreeAgent pattern: a stacking credit both sides earn after the referee's first payment, never during a trial [verified terms; inference on the choice]. Coach affiliates are a post-GA rail at the 20–30%-for-12-months norm; "show this to your EHO" has no precedent of an EHO recommending software, and printed inserts have no evidence at all.
Scheme (own page read)
Mechanic
Qualifying event
Notes
Grade
FreeAgent (UK, ~£19–33/mo)
Both sides 10% off, stacking per referral until the referrer's plan is free, then "no further benefit"
Referee's first payment; "discounts don't apply during free trials"; referrer must be paying
T&Cs updated 3 Feb 2026
[verified]
Xero UK
30% of the referee's subscription for up to 12 months, on PartnerStack
Paid plan
The "£50 after 3 months" variant page 404s
[verified; £50 reported]
Tide
£100 each
Referee spends £500 on the card within 3 months; paid within 8 weeks
Cash needs the spend gate
[verified]
Octopus Energy
£50 each; £75 each for a business account
First Direct Debit taken
—
[verified]
Metricool
Every user is an affiliate: 25% of subscriptions up to $100 per user; $200 minimum payout
Paid subscription
—
[verified]
CakeFlix
30% of tutorials and of the initial membership, 20% Master, 50% for guest tutors' own tutorials; 30-day cookie; £50 payout floor
Sale
"brands that are like minded with… our ethos"
[verified]
Stocksmith
20% of every payment for 12 months; 90-day window; PayPal, $20 floor; brand-bidding ban; aimed at bookkeepers, educators, communities
A quarter to a third of the whole budget before one referral
[verified]
"Show this to your EHO." CompliChef sells a "secure one-time access code… read-only access to your compliance records"; FoodCore's "EHO pack" sits on its £65 tier only; SFBB on GOV.UK (5 Jun 2025) is paper — "Store all your completed diary pages safely until your next visit" [verified]. ProvenBatch already ships Export EHO pack and live inspector access on every active plan [repo]. No story of an EHO recommending software was found. The plausible word-of-mouth unit is the customer showing the pack to the trader next to them — which only works if the pack and matrix are shareable by link, which FoodCore's matrix is and ours is not [repo] [inference]. Printed inserts: no evidence of Avery, Brother, MUNBYN or any UK label-stock seller running partner inserts [verified absence within budget]; the only bundle found is the reverse one (Planglow's software locked to Planglow labels).
What to ship first, and why: every account gets a code and share link in Settings; the referee gets the normal 30-day trial and, when they choose a plan, their first paid month free; the referrer gets one month's credit when the referee's first payment clears; credits stack until the referrer's year is free, then stop (FreeAgent's cap); the beta six earn a month's extension per referral instead. It costs no tooling, a referral is the only social proof available before real reviews exist, and the offer belongs on the read-only banner so the archive state becomes a reactivation path — which neither FoodCore nor Bake Diary did. Second, cheap and complementary: make the EHO pack and matrix shareable by link. Not first: coach affiliates (post-GA; keep CakeFlix 30% / Stocksmith 20%×12 as the template), cash rewards, inserts, "EHO recommends".
B10. Objections and the evidence-backed answer to each
Bottom line. The strongest ammunition is primary and free: the FSA's own text says made-to-order is not PPDS but "in anticipation of an order" is, "may contain allergens" blanket statements should not be used, non-compliance is a criminal offence with an unlimited fine, and Our Food 2023 admits "a significant backlog in the number of food businesses awaiting inspection" — the honest answer to "my council never asks" [verified]. The evidence that an objection is actually voiced comes from our own beta form and in-app feedback, Planglow's customer FAQ and one public Facebook thread on Bake Diary; verbatim forum quotes for "never inspected" and "another subscription" were not captured.
B10.1 Master list (voiced evidence: BF = production beta_requests, 9 rows; IAF = in-app feedback_items; PF = Planglow FAQ, 13 Oct 2021; SCM = Sugar Cookie Marketing Facebook thread, 10 May 2025)
#
Objection
Voiced?
Answer in our voice (facts verified unless marked)
O1
"Another subscription" / price
Indirectly: every BF respondent pays £0 for labels today (Canva 5/9, Word 1, printer/supplier 2, none 1) [repo]
£9 a month, no per-label fee. The cheapest UK label subscription otherwise is Planglow at £15–20 a month plus their labels; FoodCore starts at £25 and puts the matrix and EHO pack on £65. The label engine is complete on £9. Thirty days free, no card; on day 31 nothing is billed and nothing is deleted.
O2
"I already do it in Canva / Word / Avery"
Strongly: 5 of 9 BF; "copy and pasting all my ingredients", "manually typing them out before printing" [repo]
Canva prints whatever you type. The typing is where the mistake lives and where the time goes — your words, not ours. ProvenBatch builds the ingredients line from the packs you photographed, in the order and with the emphasis the law wants, and the label changes by itself when you switch flour brand. Keep Canva for the front; the back comes from the batch.
O3
"My council never asks" / "I've never been inspected"
Not captured verbatim; Our Food 2023: "a significant backlog in the number of food businesses awaiting inspection"; 36,000 awaiting (Sept 2026) [verified]
You are probably right that nobody has looked yet; the FSA says so itself. The label is not for the inspector, it is for the customer with the allergy, who does not wait for the backlog. And when the visit comes, it comes with whatever labels you have been printing since you registered: failing to label PPDS allergens correctly is a criminal offence with a potentially unlimited fine (FSA technical guidance). Records built as you work cost nothing extra on the day.
O4
"I just write 'may contain all 14'"
By proxy: PF asks about "advisory and precautionary labelling"; our own supplier packs "say 'may contain nuts' constantly" [repo]
Two problems. "May contain" is not an ingredients list, and PPDS food must carry the name and a full ingredients list with allergens emphasised. And GOV.UK says: "You should not use general statements such as 'may contain allergens'" — use it only for a risk "which cannot be controlled", and name the allergen. Allergic customers say they simply avoid it. ProvenBatch carries the specific "may contain" from each pack you used and shows which pack it came from.
O5
"Where is my data held?" / "AI reads my receipts?"
Not voiced as distrust; one IAF item is friction with the receipt scanner [repo]
"Database, files and accounts live in the UK" (London region). Yes, an AI reads the pack photo and the receipt, and the site says so plainly; you confirm before anything is saved. The reader is Anthropic under a UK transfer addendum; the sub-processor list is published [repo D-12/D-13]. Never say "your photos never leave your phone" — the repo has already flagged it as false.
O6
"What if you disappear like Bake Diary?"
Loudly: SCM — "Two weeks notice is waaaay short. I've been using it as a base for my taxes returns"; "2 weeks to transfer everything is almost impossible"; an 8-year user in wedding season; some never got the email [verified]
The honest answer is a promise you can check, not a promise to be here forever. Your records are exportable from day one, including on a read-only account — "visible, exportable, never deleted". Recommendation [inference]: publish a wind-down commitment — a stated minimum notice (Bake Diary gave 14 days; 90 is the credible number) and a stated export format — on the pricing page. It is exactly what that thread was asking for.
O7
"I only sell a few things"
BF: 1 of 9 "under 5" products a week, 5 "5 to 20" [repo]
Then you are the £9 tier, and the label engine is not cut down there. A butcher who "makes burgers or sausages which are prepacked to be sold on the same premises" is the FSA's own PPDS example (¶91) — three products is enough to be in scope.
O8
"I don't sell PPDS, I sell to order"
Repeatedly, to Planglow: "What about plated meals & made to order?"; cakes "bagged up when the customers choose" [verified PF]
The FSA is clear and we agree with it: "Food placed into packaging after a consumer orders it… is not PPDS" (¶93). Sold by phone or online: no ingredients list required, but allergen information "before they buy… and at the moment of delivery" (¶98–99). But most "to order" businesses also have a shed or honesty box, a market table, a cabinet, or a bake made "in anticipation of an order" — every one of those is PPDS: "packed on the premises… in anticipation of an order… includes food the consumer self selects from a chiller cabinet" (¶92); "Foods packaged and then sold elsewhere by the same operator at a market stall" (GOV.UK). Cake sheds are PPDS by definition.
O9
"Made to order so not prepacked"
Same as O8
Same answer; the test is when it went in the packaging relative to the sale, not whether you baked it this morning — "whether packaging occurs before or after customer selection, not whether it's prepared the same day" (NCASS).
O10
"Labels don't fit my packaging" / printer cost
BF: 7 of 9 own a printer (roll 5, sheets 2); the pain is typing [repo]
Use the printer you have: sheet templates and custom sizes, and the app will not print a label that has been cut off [repo]. No per-label fee. Compare CompliChef's £310 printer plus £30 a month for 500 labels, and Planglow's stock-only rule.
O11
"I'm not tech-savvy"
Voiced in beta: "this doesnt make any sense to me??" (1 Sep 2026, IAF) [repo]
Do not claim it is easy; the beta shows it is not always. Say what is true: you photograph packs and receipts instead of typing; the app reads them back to you; and the founder answers feedback directly — those items were closed with fixes within days. "Try a real recipe in the first ten minutes; if it does not make sense, tell us in the app."
O12
"My wholesaler gives me the specs" / "I use a Bako or Erudus sheet"
Not voiced by the cohort (retail buyers). Structural: Bako publishes specs; Erudus is £10–20/mo for caterers [verified]; a chocolatier in BF: brands change "frequently due to local stock issues" [repo]
A spec is the input, not the label. The spec describes the pack as sold; your label describes the batch you made from it, in order, with allergens emphasised, and it must change when the wholesaler substitutes a brand. Photograph the spec or the pack and ProvenBatch turns it into the ingredient record. If you already pay £10–20 a month for Erudus, the £9 tier costs less than the data you buy.
O13
Butchers/delis: "We use a Bizerba / Avery Berkel scale"
Not verified (vendor pages 404/429) [inference]
A label-printing scale prints the text someone typed into its product memory, which goes stale when a seasoning mix changes. ProvenBatch is where that text comes from — build the declaration from the packs, copy or export it into the scale, keep the batch record the scale does not. (No integration claim — that would be an invention.)
O14
Cafés: "We use Planglow / our EPOS prints labels"
Structural: £15–20/mo, Planglow stock only, account after a box of labels [verified]
If you are happy paying that and buying their stock, it does the sandwich job. ProvenBatch is £9, prints on any stock, and the label is built from the packs in that batch and costs the recipe against the receipt at the same time. The switch cost is real; the receipt and pack camera is how it is paid.
O15
Stalls: "I label by hand on the day"
Structural: NCASS — "if you… pack your own food product before it is consumed… you are affected by the PPDS rules"; FSA ¶89(iii) names "market stalls, mobile sales vehicles" [verified]
Hand-written is not illegal; incomplete is. A hand-written label still needs the full ingredients list with every allergen emphasised, every product, every week. Print them the night before from the batch you actually made; if you swapped brands, the label already changed. Anything wrapped after the customer chooses needs no label — you just need to be able to say what is in it (¶93).
O16
"I'm a home business, too small for this"
Council pages: home bakers "must register with us at least 28 days before opening", free (Runnymede) [verified]
Registration is free and cannot be refused; PPDS applies from the first packed cake; the same £9 does the label and the records the council will eventually ask for.
O17
"You have no reviews"
True today
Say so. "We are new; six businesses have used it since 1 September and their feedback is in the release notes." Point at the changelog and the export and read-only terms, not at praise.
B10.2 The ten most likely objections per vertical (mapped to the master list; the note is the vertical's twist)
Home bakers / cake sheds: O2, O1, O3, O8/O9, O4, O6, O11, O10, O16, O17. Twist: lead with the shed — self-select from a cabinet is PPDS by the FSA's own example.
Caterers / meal-prep: O8 (most orders are distance sales: no ingredients list required but allergen info before purchase and at delivery), O12, O1 (they may already pay FoodCore or CompliChef for food safety), O14, O4, O3 (they are inspected more; keep the answer to the offence, not the backlog), O5, O13-style "our kitchen system does it", O11 (staff will use it), O7. Twist: the meal-prep pot in the fridge for collection was "packed in anticipation of an order".
Butchers: O13, O7 (the FSA's own butcher example), O12 (seasonings, rusk), O3, O1, O10 (scale rolls), O11, O5, O4 (shared mincers — the one place PAL is justified, and it must be named), O15. Twist: the seasoning mix is the recurring substitution risk.
Delis / farm shops: O8 ("we sell other people's products" — prepacked by another business is not PPDS, ¶95, but your scotch eggs and salads are), O13, O12, O1, O3, O7, O4, O10 (jars, tubs), O5, O11. Twist: split the shelf — bought-in prepacked vs made-and-packed-here.
Cafés / sandwiches: O14, O9 ("made in front of the customer" is genuinely not PPDS; the grab-and-go fridge is), O1, O10, O11, O3, O4, O12, O5, O7. Twist: GOV.UK's "burger under a hot lamp" example.
Market stalls / street food: O15, O8, O1 (NCASS/NMTF members already pay memberships), O10 (weather), O3 (a different council every market), O4, O11 (phone-only), O6, O7, O16. Twist: packed at home, sold from "a moveable and/or temporary premises" by the same business is PPDS (¶89(iii)).
Chocolatiers: O12 (brand substitution — the chocolatier in our pipeline typed it), O1, O4 (nuts and milk — where specific PAL is real), O10 (tiny labels; note the <10 cm² rule, ¶96), O8 (gifting and online), O3, O2 (Canva branding), O7, O11, O17. Twist: per-batch pack tracking is exactly the problem they typed into the form.
Preserves / sauces: O8 (online sales), O7 ("I make one jam"), O1, O10 (jar labels), O2, O4, O3, O16 (registering across councils for markets), O12 (pectin, sugar), O6. Twist: jars sold at a market by the maker are PPDS; jars supplied to a farm shop are prepacked and need the full label, which the app also handles [repo].
C. The plan
C11. The 90-day acquisition plan — one page, 1 October to 31 December 2026, one founder, ~10 h/week, £150/month
Everything here is [inference] built on §A–§B. The 1 October hook is the fifth anniversary of Natasha's Law commencing; every pitch below leads with it.
Week 0 (before 1 Oct) — measurement and the four assets. Carry UTM/gclid into the signup record; add "how did you hear about us?"; one promo code per partner. Ship the one-product first session and the day-1 concierge email. Publish: "Bake Diary alternative UK", "ProvenBatch vs FoodCore" (facts only), the free no-signup PPDS label checker, and "After you register: what the inspector will ask to see". Claim the directory listings for the badges. Referral code in Settings scaffolded.
Priority
Channel
First concrete action (date)
Weekly measure (Fridays)
Kill / cut criterion
1
Activation + concierge
1 Oct: day-1 "we'll do your first product with you" live; founder does the concierge personally, capped at 5/week
% trials with a real label from a photographed pack by day 7; concierge accepted vs declined and each cohort's conversion
Week 4: <15% activated → pause channels 3–4, spend their hours here. Never kill; this is the funnel
2
Coach and admin partnerships, podcasts
1–3 Oct: write to Kate Tynan (hello@ on her site) and Annie Bennett (contact form) offering a free guest labelling session before the Christmas order run + 30% recurring affiliate; pitch The Business of Cake Making via its pitch page; British Sugarcraft Guild for Sugarcraft News
Sessions agreed; attributed trials per session via promo code
Week 4: no session agreed → one follow-up a month, hours to priority 3
3
Exact-match search + the four pages
1 Oct: Google exact-match on "allergen label software", "natasha's law software", "ppds label software", "bake diary alternative", £3–5/day; pages live
Exact-match impressions (<300/week = no volume); cost per started trial at £100 spent
>£40/trial after £100 → kill paid, keep pages. Honest target ≤£8/trial given §0.7
4
Metricool social + reactive routine
Queue runs; 1 Oct anniversary post; a post within a day of any council enforcement release, ingredient recall or TikTok Shop/eBay policy change; £25 boost of the best launch Reel in week 2
Link clicks to /trial (≥20/week by week 4); ≥1 attributed trial/week
Missed at week 4 → Autolist recycling + reactive posts only
5
Institutional rails
Week 1: Protectivity guest article "Natasha's Law at five: what 'labelled accurately' means" (contact@); Erudus integration-partner enquiry (support@); NMTF member-deal call (01226 749021); Bucks Council food-safety team offered a free home-food-business guidance page (food.safety@); Mission Kitchen perk + talk (hello@)
Replies; one live listing/article by 30 Nov
Nothing live by day 60 → keep only Erudus and Protectivity warm
6
Referral credit
Weeks 3–4: "give a month, get a month" live, on the read-only banner; beta six get a month's extension per referral
Codes used per 100 active accounts; read-only reactivations at 30/60 days
None (one-off build); review monthly
7
Cake International visit
6–8 Nov, NEC: OLBAA stand C40, BSG demos, coaches; trial cards only where invited
Follow-ups generated
n/a
Monthly gates. 31 Oct: activation ≥25% or everything else pauses; at least one partner session dated. 30 Nov: cost per started trial known for search and social; one institutional listing live; referral shipped. 31 Dec: trial-to-paid on the October cohort (day-30 aged) — plan on 4–8%, target ≥10%; the decision on whether the £6–20-per-trial ceiling survives is made on this number, not before.
Not in the plan, by decision: daily manual group posting; cold email or DMs to anyone not a verified Ltd/LLP; exhibiting; any list purchase; Reddit, forums, LinkedIn; testimonials before real ones exist.
C12. The twenty leads to pursue first
All UK, all with 2025–2026 own-site evidence read on 20 Sep 2026 unless marked [reported]. Route column applies §B7: a published business address at a Ltd company, association, public body or publisher is a corporate subscriber; a coach or small operator whose incorporation is unverified gets their own contact form or explicit pitch page, never a cold email. Ranked by trust × reach × ease × timing.
#
Lead
Type
Why (own-page evidence)
The ask
Lawful route
Timing
URL
1
Kate Tynan — The Cake Business Club (+ free FB group A Bigger Slice)
Coach / admin
"private community of 100+ cake makers"; "coached 550+ since 2022"; Coach of the Year 2025 [verified]; group size not verified
Guest "allergen labelling that proves itself" session; named tool in the Foundations set-up module; 30% recurring affiliate
hello@thecakebusinessclub.co.uk is published for business enquiries; incorporation unverified → write as a one-to-one partnership enquiry
Now; second touch after Cake International
https://www.thecakebusinessclub.co.uk/
2
Annie Bennett — The Profitable Baker Academy (+ The Home Baking Business Community UK)
£150/yr membership incl. liability cover; "actively seeks partnerships"; /deals lists a bank, the AA, solicitors, insurer, accountant — no software [verified]
Member deal on /deals; piece in the magazine/e-bulletin
Phone 01226 749021 (no email in fetch; /contact 404) — corporate body
Oct–Nov, before Christmas markets
https://www.nmtf.co.uk/deals
6
The Business of Cake Making Podcast (Daisy Cake & Co)
Podcast
"85.7K downloads, 186 episodes"; Ep 162 covered Bake Diary; co-host appeal since Jun 2026 — format in transition [verified]
Guest/co-host slot: "the label that proves the batch"
Public pitch-an-episode page (explicit invitation)
Now; ask whether autumn episodes resume
https://thebusinessofcakemaking.podbean.com/
7
Encore Kitchens (ex-Foodstars)
Kitchen landlord
"25 UK locations", "550+ Kitchens", "300+ Brand Partners", Members Perks page [verified]
Members Perks listing; onboarding leaflet
Book-a-tour / contact form (corporate)
Now
https://www.encorekitchens.co.uk/
8
The Makers Market (NW & Midlands)
Market operator
"25+ markets per month"; T&Cs require allergen information displayed; named organiser email [verified]
Trader-onboarding resource "your labels for market day"; newsletter mention
roseann@themakersmarket.co.uk — a business enquiry about a maker resource, not a marketing send
Late Sept–Oct (Christmas trader intake)
https://www.themakersmarket.co.uk/
9
Farm Retail Association
Association
Farm Supplier Membership tier; Conference 2026; ~220 members in the finder [verified]; supplier price not shown
Supplier membership → directory listing + member discount; a PPDS piece for farm-shop counters
"Get in Touch" form / 01423 546214 (corporate)
Before the 2026 conference
https://farmretail.co.uk/
10
Buckinghamshire Council — Food Safety team
Council / EHO
Registration page with team email; no home-food-business guidance page yet [verified]; precedents: West Norfolk's cake-maker PDFs, Wigan's home-baker pack
Offer a free "starting a home food business — labels and records" page in the West Norfolk style, tool named once as one option; EHO feedback on our records view
food.safety@buckinghamshire.gov.uk (public body — corporate)
Next five, verified but ranked out for cost or timing: Craft Bakers Association (Industry Supporter packages from £750+VAT; a £200 e-shot to members is the cheap test; Jan 2026 PDF); Scottish Bakers (Allied Trade membership, info@scottishbakers.org, Q1 2027); Dead Simple Accounting ("Accountant for Caterers & Bakers", publishes a tools list — one email); Enterprise Nation (adviser profile £20/mo; partnerships mailbox); Bako (five depots, "In the Mix" newsletter; a pack-data feed in 2027). Excluded as unverifiable or conflicted: NCASS (bundles its own Digital Food Safety System); Pilgrim Foodservice (sells Planglow); Guild of Jam and Preserve Makers (domain redirects elsewhere); Karma Kitchen (domain for sale); Scottish Craft Butchers (DNS); Finch Bakery (a retail brand, not a coach); Reflex Labels (industrial). Anaphylaxis UK and NARF are active and high-trust, but a for-profit's approach to a bereaved family's charity must be a donation or Natasha's Day support first, never a marketing ask — parked until there is revenue.
Decisions worth reopening
None of the §4 decisions is contradicted by anything found in this pass, so none is proposed for reopening. The one number that does need changing is not a §4 decision but a §3 planning assumption: the "£6–20 per started trial" ceiling only holds at ~20% trial-to-paid, and the best opt-in, no-card benchmark is 4–8% (§0.7, §B8). Two adjacent points, stated so they are not mistaken for reopening: addressed post to a business at its published shop address and email to a verified limited company are lawful without consent under the current ICO guidance (§B7) and are not excluded by §4's "no cold email or DMs to sole traders"; they are low-volume, manual routes for delis, farm shops and cafés, not channels.
What I could not find
Verbatim customer voice. No public Facebook posts, Mumsnet threads, TikTok captions or YouTube comments could be read (403/404/CAPTCHA on every attempt); "takes me hours", "my EHO said", "may contain all 14", "never been inspected", "another subscription" were not captured from customers. The substitutes are our beta form, in-app feedback, the FSA's 2019 consultation, Planglow's customer FAQ, review-site quotes and the FSS consumer research.
The FSA's PPDS evaluation (April 2023) full chapters — food.gov.uk returns 410; figures rest on the FSA blog, Food Manufacture and Roythornes. Later Food and You 2 hypersensitivity chapters likewise; only the 2021 channel-confidence figures are verified.
National statistics on PPDS enforcement since Oct 2021 — none published; the FSA says England/NI food standards data is "currently unavailable". No prosecution of a home baker or market stall for PPDS labelling was found.
Any UK trial-to-paid figure for SMB SaaS; any controlled measurement of read-only vs lock-out, or of a done-for-you setup offer, at a sub-£40 price. The "25%/60%" and "72 hours" statistics have no primary source.
Paid-ad activity of any analogue (Google Ads Transparency is a JavaScript shell; Meta's Ad Library 403). Follower counts for FoodCore, Nutritics, Bakesy. Any founder interview for FoodCore, AllergenKit, Bakesy or Bake Diary; Bake Diary's founder name and stated closure reason.
Council or FSA pages linking third-party tools — none in ten pages; the East Suffolk PDF (405) may be the only instance beside West Norfolk.
Blocked sites: Pilgrim Foodservice (403; its Planglow arrangement is confirmed from Planglow's side), Avery UK, Brother's Natasha's Law blog, British Baker / bakeryinfo (405), The Grocer (405), Etsy's food policy, Amazon UK seller help, alerts.food.gov.uk, TPS Online (empty render), MPS licence costs, WhatDoTheyKnow, Guild of Fine Food, Q Guild, Bury Market, KFMA, NMTF's Market Near Me (522), FSB member benefits (JS), Bizerba and Avery Berkel, Kit's concierge page, Canva's affiliate page, Bird & Bird's "free deals as a sale" article (paywall).
Not checked (search budget): Cake Shed, Jam Jar, Homemade.io, Orderly, Shoppable and Shopify/Square/SumUp seller directories; Cake Craft Company (verified active only), BakeryBits, Kite, Reflex, Label Bar, Tommy Labels, Sticker Mule, Dymo; Kitchen Table Projects, Cook Space; BFP, Costco, Booker's and Brakes's small-business Natasha's Law marketing (Brakes and KFF PDFs downloaded but no extractor); Levy Market, Geraud, Market Place Europe; Business of Cake, Cake Business School; Xero/FreeAgent partner directories; Monzo Business referral; coach graduate showcases; Etsy's and Amazon's current policy wording; whether TikTok Shop's invite-only food gate pre-dates June 2025.
Member counts for A Bigger Slice, The Home Baking Business Community UK (7,000+ is an extract), FRA, BSG, CBA; Farm Retail Association supplier-membership price; NMTF partnerships email; OLBAA's £50 sponsorship; Cake International's visitor count; CMTIA's and BBF's emails.
Pending ICO material: the "PECR advice for small organisations" update (Drafting, due Summer 2026); the promised guidance on the higher PECR fines; TPS/CTPS licence fees.
One open thread: Safer Food Group's allergy course cites "UK Government statutory guidance published July 2026" — not identified; if real, it is a launch-week content hook.
The FHRS figures are a 14-authority sample (28,163 records); the ~95–100k private-address and ~55k awaiting-inspection totals are extrapolations, not counts.
Sources
All accessed 20 Sep 2026. R = fetched and read (or pulled and parsed); E = extract or secondary only. Grouped by the pass that used them; a source used by several passes is listed once.
Official — FSA, FSS, GOV.UK, legislation.gov.uk
FHRS open data — https://ratings.food.gov.uk/open-data (dataset 17 Sep 2026) R · API help — https://api.ratings.food.gov.uk/help (undated) R · /BusinessTypes/basic, /Establishments, /Ratings, /SortOptions, /Authorities/basic (live, 20 Sep 2026) R · per-LA files https://ratings.food.gov.uk/api/open-data-files/FHRS{code}en-GB.xml (17 authorities, extracts 16–20 Sep 2026) R · FSA data catalogue entry — https://data.food.gov.uk/catalog/datasets/38dd8d6a-5ab1-4f50-b753-ab33288e3200 (undated) R · Open Government Licence v3 — https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/ R
FHRS guidance for businesses — https://www.gov.uk/government/publications/food-hygiene-rating-scheme-fhrs-guidance-for-businesses/food-hygiene-rating-scheme-fhrs-guidance-for-businesses (updated 30 Jun 2026) R · Register a food business — https://www.gov.uk/food-business-registration (25 Jun 2026) R · register.food.gov.uk privacy notice (undated) R and /new landing R · Starting a food business — https://www.gov.uk/guidance/starting-a-food-business (updated 19 Aug 2026) R · Running a food business — https://www.gov.uk/running-food-business R
FSA September 2026 Board: Local authority performance update — https://www.gov.uk/government/publications/food-standards-agency-september-2026-board-meeting/local-authority-performance-update-food-standards-agency-september-2026-board-meeting R · Future of Food Regulation report — …/future-of-food-regulation-report-to-the-fsa-board-september-2026 (3 Sep 2026) R · Chief Executive's report R · Business Committee report (15 Sep 2026) R
FSA December 2025 Board: position on the Codex PAL standard — https://www.gov.uk/government/publications/fsa-25-12-05-fsa-position-on-the-codex-precautionary-allergen-labelling-standard-including-allergen-thresholds/… R
Audit of allergen controls — Rhondda Cynon Taf (9–11 Feb 2026) — https://www.gov.uk/government/publications/audit-of-allergen-controls-and-relevant-open-audit-actions-rhondda-cynon-taf/… R
PPDS guides: introduction (updated 17 Jul 2026) — https://www.gov.uk/government/publications/introduction-to-allergen-labelling-for-ppds-food/… R · bakers (9 Jul 2021) R · butchers, mobile sellers and street food, event caterers, restaurants/cafés/pubs (all 17 Jul 2026) R · labelling guidance for PPDS (17 Jun 2021) R · Allergen guidance for food businesses (updated 17 Jul 2026) R · technical guidance (23 Aug 2023, updated Mar 2025) R · PAL page — https://www.gov.uk/food-labelling-and-packaging/precautionary-allergen-labelling-pal R · allergen labelling for manufacturers (14 Dec 2017) R · FSA technical guidance June 2020 Part 3 ¶86–103 (PDF mirror at abcfoodlaw.co.uk) R
Food Information Regulations 2014 regs 10, 11 — https://www.legislation.gov.uk/uksi/2014/1855/regulation/10 R · LASPO 2012 s85 R · Food Safety Act 1990 ss14, 35 R · Regulation 1169/2011 Art 44 R · Impact Assessment 2019/144 (PDF) — https://www.legislation.gov.uk/ukia/2019/144/pdfs/ukia20190144en.pdf R · 2019 consultation outcome (25 Jun 2019) — https://www.gov.uk/government/consultations/food-labelling-changing-food-allergen-information-laws/outcome/summary-of-responses-and-government-response R
FSA blog, evaluation of Natasha's Law (19 Jul 2023) — https://food.blog.gov.uk/2023/07/19/… R · Surveillance Sampling 2022-23 (22 Feb 2024) — https://science.food.gov.uk/article/127614-… R · Retail Surveillance 2024/25 (26 Jun 2025) — https://science.food.gov.uk/article/140609-… R · Our Food 2023 (PDF, 8 Oct 2024) R · Our Food 2024 (PDF, 19 Jun 2025) R · Annual incidents report 2024/25 (25 Jun 2026) R · Allergen information for non-prepacked foods (7 Oct 2025) — https://science.food.gov.uk/article/142304-… R · SME out-of-home research (3 Oct 2024) — https://science.food.gov.uk/article/123212-… R · Food and You 2 Wave 11 (31 Mar 2026) R · Ipsos, Food and You 2 Wave 2 (10 Aug 2021) R · UK Food Security Report 2024 Theme 5 (11 Dec 2024) R · FSA press release via WiredGov (24 Feb 2022) R · Mustard/peanut guidance (30 Sep 2024) R · Allergy alerts sign-up pages R · School food standards allergy guide (14 Sep 2026) R · SFBB on GOV.UK (5 Jun 2025) R
FSS PPDS consumer research (PDF, June 2023) — https://www.foodstandards.gov.scot/sites/default/files/migration/downloads/PPDSconsumerresearch_report.pdf R
ICO: Direct marketing guidance hub and sub-pages (updated 28 Apr 2026) — https://ico.org.uk/for-organisations/direct-marketing-and-privacy-and-electronic-communications/direct-marketing-guidance/ R · electronic mail guidance (28 Apr 2026) R · live calls guidance (under review) R · Guide to PECR pages (under review) R · plans for new guidance R · DUAA summary (19 Jun 2025) and "what does it mean" (19 Jun 2026) R · statement on commencement (5 Feb 2026) R · charities soft opt-in news (28 Apr 2026) R · "One year on" (23 Jun 2026) R · legitimate interests (23 Mar 2026) R · consent guidance R · right to be informed R · B2B marketing page R
PECR 2003 regs 2, 19, 21, 22 (as at 5 Feb 2026), 23, 26 — https://www.legislation.gov.uk/uksi/2003/2426/… R · UK GDPR Art 6 and Annex 1 R · DUAA 2025 ss110, 115, Sch 13 R · SI 2025/904 and SI 2026/82 R · Clifford Chance (6 Feb 2026) R · MPS — https://www.mpsonline.org.uk/ R (thin) · TPS Online E (empty render)
Councils: Salford, Ealing, North Norfolk, Oldham (registration; prosecution 18 Dec 2023), Worcestershire RS, Brighton & Hove, North East Lincolnshire (1 Mar 2022), Haringey, Basingstoke & Deane, Luton, Runnymede, North Yorkshire, Scottish Borders (25 Jun 2021), Reigate & Banstead, Wigan (PDF v2 14 Apr 2021), King's Lynn & West Norfolk, Buckinghamshire, Newcastle (9 May 2025), Pembrokeshire (30 Sep 2025), Barnsley (27 Jan 2026), West Sussex (27 Feb 2026), Hartlepool (4 Mar 2026), South Norfolk & Broadland (7 Oct 2022) — all R · Cornwall food register on data.gov.uk (29 Sep 2020) R · Kent & Medway Growth Hub (10 Feb 2022) R · WhatDoTheyKnow E (403) · East Suffolk PDF E (405)
Charities, trade bodies, press
Anaphylaxis UK: Celia Marsh PFD (7 Dec 2022), mustard incident (21 Oct 2024), bake-sale guidance (7 Aug 2025), PAL update (4 Sep 2023) R · NARF: Owen Carey inquest, National Allergy Strategy (22 Apr 2026), TikTok coverage (3 Jun 2025), What is Natasha's Law R · Allergy UK Awareness Week R · CIEH press release (6 Jun 2025), EHN (3 Jul 2025), four-step plan (13 Oct 2021) R · Food Manufacture (4 Jun 2025; 2 Oct 2023; 10 Apr 2019) R · Roythornes (2023) R · C&C Solicitors (Mar 2026) R · Walker Morris (3 Oct 2023) R · Business Companion PPDS (Mar 2025) and bread/cakes (Jun 2025) R · A Little Bit of Stone (15 Sep 2026; 11 May 2026) R · Magic Sparkles (13 Sep 2026) R · Labelservice blog (16 Sep 2026) R (claim contradicted) · GeraEats recalls (25 Jun 2026) E · Speciality Food Magazine (advertise page; 1 Oct 2021 article) R · Exhibition News (Cake International) E · British Baker / bakeryinfo E (405) · Google News RSS listings E
Platforms: TikTok Shop UK F&B listing guidelines (updated 25 Aug 2026) — https://seller-uk.tiktok.com/university/essay?knowledge_id=7753844604159745 R · eBay UK food policy R · Meta commerce policies R · Skilltopia (29 Jan 2026) R · High Speed Training on Etsy (23 Jul 2021) R · Etsy policy E (403)
Referrers, partners, leads (own pages)
Farm Retail Association (finder, membership, home) R · NMTF (home, /deals, /details, /faq) R · NABMA R · Pedddle R · farmers-market.org.uk R · Stallfinder R · NCASS (home, find-a-caterer, PPDS page, Karma Kitchen pages) R · Nextdoor R · National Craft Butchers R · Q Guild E · Scottish Bakers R · Craft Bakers Association (packages PDF, Jan 2026) R partial · British Sugarcraft Guild R · The Makers Market (home, T&Cs) R · Mission Kitchen R · Encore Kitchens R · Karma Kitchen (for-sale page) R · Broadland Food Innovation Centre R
High Speed Training (Paid On Results; hub; Savvy Baker case study 21 Aug 2026) R · Virtual College (Awin affiliates) R · iHASCO and Food Alert R · Smart Horizons R · Safer Food Group R
Protectivity (product page 7 Jan 2026; FAQ; knowledge centre; contact) R · CMTIA R · Superscript R · PolicyBee R · Simply Business (guide updated 31 Mar 2026; US partners page) R · Becca's Bouqcakes blog R
Dead Simple Accounting R · Livingstones Accountants R
Brother UK PPDS page R (via curl) · MUNBYN UK R (terms blocked) · NIIMBOT UK R · Avery UK E (403) · Handy Labels E (403) · Good Cake Day R · Cake Craft Company R · OLBAA (uk, us podcast) R; sponsorship price E · Reflex Labels R
Planglow (Pilgrim article 7 Dec 2021; Booker 11 Apr 2024; pricing; trial form; software page; FAQs 13 Oct 2021; 2021 guide) R · Pilgrim Foodservice E (403) · Bidfood Natasha's Law R · Craft Guild of Chefs (21 Jan 2022) R · Erudus (home, caterers, retailers, wholesalers, integration partners, checklist 20 Sep 2022) R · Bako R · Brakes and KFF PDFs E (no extractor)
CakeFlix affiliate pages R · Annie Bennett (links, join, contact) R · The Cake Business Club R · A Bigger Slice (login wall) E · The Business of Cake Making (Podbean; Daisy Cake & Co) R · Slice of Success RSS (13 May 2025) R · Janelle Copeland (12 May 2026) R · Sugar Cookie Marketing Facebook post (10 May 2025) R
Enterprise Nation R · BIPC via Grow London Local R · Buckinghamshire Business First R · FSB E (JS) · Cake International (dates, exhibitor list) R · Anaphylaxis UK / NARF (activity) R
Analogues and trial benchmarks
FoodCore (home, pricing, about, product, blog, vertical and FAQ pages; Trustpilot; GetApp) R · AllergenKit (home, about, matrix, Natasha's Law) R · CompliChef (.co.uk home, labels, switch; Trustpilot; .com checked) R · Nutritics (home, pricing; Capterra; GetApp; Trustpilot; restauranttools.ai) R · Kafoodle (home, trial request; Capterra; GetApp; Trustpilot) R · Bakesy (home, pricing, resources; App Store GB/US; iTunes lookup API; Google Play) R · Stocksmith (home, pricing, affiliates, compare/bake-diary, blog 31 Mar and 1 Jul 2026) R · Bake Diary (domain check; iTunes/Play search) R; €6.95 E · BakeMargin (20 Feb 2026) R · BakerInbox (15 Jun 2026; pricing) R · Bake Boost R · Crumb Coach (18 Aug 2026) R · Paddl R · BakeProfit R · CakeBoss (+ Capterra) R · Allergen Checker R · Marka R · Castiron E · Cakenote E · Baking It E (403) · Google Ads Transparency (JS shell) and Meta Ad Library (403) E
ChartMogul/Poyar 2026 conversion report (growthunhinged.com; chartmogul.com reports 1 and 2; Poyar note 4 Feb 2026) R · ChartMogul GTM report 2025 R · Lenny's Newsletter (1 Aug 2023; 25 Oct 2022) R · First Page Sage (5 Sep 2025) R · Ada Chen Rekhi compilation R · Sixteen Ventures (14 Jun 2014) R · Recurly (16 Jan 2025; blog) R · Yoganarasimhan et al. (arXiv 2006.13420) R · Zhang & Duan 2025 (PMC12217587) R · Startups for the Rest of Us ep. 758 (18 Feb 2025) R · Userpilot (two posts, 2026) R secondary · Appcues (4 and 12 Jun 2026) R · Userlist (updated 22 Apr 2026) R · earlystagefounder.com (Groove) R secondary · Mixergy (Groove) R · Failory (ConvertKit) R secondary · Podia /switch R · Poyar reverse trials (29 Jun 2022) R · Basecamp pricing and help R · Todoist help R · Notion help R · HEY FAQs R · Databox (12 Apr 2023) R · Crazy Egg, UserGuiding, Chameleon, 1Capture, Ordway, shno.co R (checked for unsourced claims)
Referral: FreeAgent T&Cs (3 Feb 2026) R · Xero UK refer-a-friend R · Tide R · Octopus R · SumUp Pay (6 Mar 2026) R · Starling R · Squarespace Circle R · Shopify Affiliates R · Notion Affiliates R · Metricool affiliate R · Canva via linkjolt E · Rewardful, FirstPromoter, Tolt, ReferralCandy, PartnerStack pricing R
Repository and production data [repo]
outputs/gtm/beta-pipeline.md; production beta_requests (9 completed rows) and feedback_items (10 most recent), read 20 Sep 2026, no personal data quoted · outputs/docs/gtm-channels-research.md · outputs/docs/metricool-growth-research.md · master competitor sheet · outputs/RESEARCH.md · outputs/docs/go-to-market-decisions.md (D-12/D-13) · outputs/USER-GUIDE.md · outputs/gtm/drafts/feedback-for-free-year-definition-2026-09-14.md · outputs/docs/provenbatch-website-content-foundation.md
Market
Last updated 15 Sep 2026
Metricool as the growth engine for the 1 October launch — tiers, ads, and a loop a robot can run
Researched 15 Sep 2026 against app v0.381.0 on production, the live Metricool brand (ProvenBatch, id 6676783, read through the claude.ai Metricool connector that morning), Metricool's own pricing and help-centre pages, xAI's connector and automation pages, and UK ad-cost benchmarks published in 2026 — none of it from memory (AGENTS.md #19b). Where a page could not be fetched it is marked below rather than filled in. Deliverable: this document plus four one-page decision memos in outputs/gtm/drafts/ dated 15 Sep 2026. RESEARCH.md §28 carries the headlines. Nothing here was built, bought, scheduled or spent: every action needs Dave's word (AGENTS.md #10 shape).
✅ Status update 15 Sep 2026 (Dave): brand ProvenBatch (6676783) is now on Starter monthly. The Free 20 posts/month wall no longer applies. Historical Free-plan notes below remain as of the morning research read; treat Starter as current for ops.
The question, as Dave put it (15 Sep): how do we maximise Metricool to grow ProvenBatch fast from the 1 October public free-trial launch, as autonomously as possible, with Grok orchestrating, considering every Metricool tier, products and services we are not using yet, and ads — wherever ROI can be forecast. Dave's framing for the answers: paid channels are fully reopened (the 14 Sep Google-only decision is on the table), the forecast plans against ~£150 a month all-in for the first 90 days, and weekly batch approval is the ceiling on autonomy.
0. The answer in one page
Stay on Metricool Free until a specific wall is hit, then go straight to Starter (~£14–17/mo). Of the four tiers, Starter buys the only three things we would actually use — unlimited posts (Free is 20/month; the queue is already running at ~15 a fortnight), unlimited analytics history (Free keeps 30 days, which means the launch baseline is lost by November), and Flows once it becomes a paid add-on. Advanced (~£37–46/mo) adds the post-approval workflow, Looker Studio, API and Metricool's hosted MCP — none of which we need, because the claude.ai Metricool connector already works on the Free plan (verified 15 Sep: brand settings, scheduled posts, analytics, best-time all returned; only the review-mode create returns 403). §2.
Nothing can be judged on ROI until two measurement gaps close, both cheap and cookieless: the Cloudflare Web Analytics toggle (D-27 already sanctions it) and carrying the utm_*/gclid parameters the site already forwards into the app's signup record. Today the site sees only a click that leaves for the app; the app sees a business with no origin. That gap makes every channel below worth exactly £0 on paper. §4.
Paid, reopened and modelled: Google exact-match Search stays the first bet, a £25 Meta boost of the best launch-week Reel is the second, and a one-week TikTok Spark Ads test is the third — all inside £150/month, in that order, each with a cost-per-started-trial gate. The economics have not changed since 31 Jul: allowable CAC £30–100, so at a 10–20% opt-in trial→paid rate the most we can pay for a started trial is roughly £6–£20. Anything over £40 after £100 spent is killed. Meta conversion campaigns and anything needing a pixel stay off until the consent question (#1822) is decided. §5–§6.
The autonomous loop is one Grok automation firing twice a week against one approval card: Monday it drafts the coming week (posts, ≤2 trend ideas, one boost candidate, last week's numbers, ads against their kill rules) and delivers a single card in Charlie/Grok chat (the delivery surface since 19 Sep 2026); Dave's ✅ / ✏️ / ❌; Tuesday's fire schedules the approved batch live in Metricool. Metricool's own automations (Autolists for evergreen recycling, best-time slots, a Flows keyword that DMs the trial link) do the rest. What it cannot do yet: Grok holds no Metricool connector. Adding Metricool's hosted MCP as a Grok custom connector is a ten-minute test; if it 403s on Free, that is the Advanced trigger, or the existing Claude growth routine stays the Metricool executor. §7.
The repo is stale on Metricool in two places, corrected today: the queue is not drafts (15 posts 15–27 Sep, autoPublish: true, draft: false), and the brand carries three connected ad accounts — Facebook Ads act_1804831247189540 (the repo names a different id), Google Ads 9992474342, and TikTok Ads 7670576612999725072 (nowhere in the repo). §1.
Ceiling check. Starter Metricool (~£17 monthly) + Google (~£60–100 actual, volume-limited) + one boost (£25) + generated b-roll clips (~£12) ≈ £115–155/month. A TikTok Spark week (£140 at the £20/day ad-group minimum) only fits by pausing Google that week. The larger options — Meta conversion campaigns at the £300–600/month test floor the July research cites — are modelled in §6 as gated escalations Dave would have to fund separately.
1. Where we are on 15 Sep 2026
1a. The live brand, read this morning
Item
Live value (15 Sep, getBrandSettings)
Repo says
Brand
ProvenBatch, id 6676783, owner social@stellaapps.com, timezone Europe/London, joined 5 Aug 2026
✅ matches social-accounts.md
Facebook Page
1298004883398020
✅
Instagram
provenbatch
✅
TikTok
provenbatch
✅
YouTube
UCEkBPr1-Yq6HMg0B09LVuBA
✅
Facebook Ads
act_1804831247189540
⚠️ social-accounts.md names act_1357764382546345
Google Ads
9992474342
✅ named in social-accounts.md
TikTok Ads
7670576612999725072
❌ not in the repo at all
Plan
Free
✅
So the ad-account plumbing the 14 Sep memos list as "need from Dave: Meta ad account" is at least partly done: all three networks' ad accounts are connected to the Metricool brand. Whether the accounts hold a payment method is not visible from the API and remains Dave's to confirm.
1b. The queue is live, not drafted
getScheduledPosts (15 Sep–31 Dec) returned 15 posts, 15–27 September, every one autoPublish: true, draft: false, created 14 Sep 12:57–12:58 by social@stellaapps.com. Nine are Instagram Reel + TikTok pairs (three of them also YouTube Shorts), four are Facebook-only text posts, four are YouTube uploads on 18 Sep (overview, walkthrough, vertical overview, longform). Every caption closes "Opens 1 October. provenbatch.co.uk/trial". One extra row dated 31 Dec is a draft placeholder reading "[DROPPED — missed 14 Sep 08:28 slot; Charlie 15 Sep. Do not publish.]".
This supersedes three repo claims at once: HANDOVER's "Posts are live as drafts; nothing publishes while drafted", the metricool-drafting skill's "review mode only — this skill never takes a post out of draft", and the 30 Aug beta-line rule (the queue now uses the 14 Sep amendment (e) pre-launch line, correctly). The HANDOVER paragraph is corrected in the same commit as this document; the skill is not edited here (documents only), but §7 says what it should say.
Nothing is scheduled after 27 September. The launch week itself — 1–7 October — is empty in Metricool. That is the first thing the loop in §7 has to fill, and it is why the loop needs to exist before 1 October rather than after.
1c. The 30-day baseline (15 Aug – 14 Sep, getAnalyticsDataByMetrics)
The numbers a launch will be measured against. Small, and worth writing down precisely because the Free plan's 30-day history means they will be unreadable after mid-October unless exported.
Network
Published
Views / impressions
Reach
Interactions
Followers (14 Sep)
TikTok
11 videos
~3,100 video views; ~205 profile views
—
14
3
Instagram
14 Reels
~540 Reel views
~490
~27
connector returned 0 — read in-app
YouTube
4 uploads
~116 video views
—
—
1–2 subscribers
Facebook Page
text posts (count not returned)
~1,600 page views
—
~34
85 (+83 in the period, two jumps of 21 and 33 on 20 and 25 Aug)
Read: TikTok is already the reach channel by an order of magnitude — one Reel/TikTok pair on 24 Aug did 816 TikTok views against 160 on Instagram; Facebook is the credibility surface the July research predicted (people look, few interact); YouTube is search inventory that compounds slowly. getBestTimeToPostByNetwork for TikTok (7–14 Sep) peaks at 10:00, 12:00 and 18:00 UK on Wednesday–Friday, with the weekend a clear trough; for Instagram it returns all zeros — the account has too few followers for Metricool to compute anything, which itself says the 19:00–21:00 slots in social-content.md §3 are a guess and TikTok's measured slots should lead.
1d. What the MCP actually does on the Free plan (empirical, 15 Sep)
❌ 403 — the "post approval system" is an Advanced feature
Metricool's own MCP marketing page says access "requires the Advanced plan or higher". The claude.ai connector evidently does not enforce that on reads or plain creates. Do not build on the assumption it stays that way — the loop in §7 keeps a fallback — but do not pay for Advanced to get something we already have.
1e. The skill contradiction, for the record
.claude/skills/metricool-drafting/SKILL.md mandates createScheduledPostForReview only; .claude/skills/social-video/SKILL.md (and the removed meme-video experiment README) say the Free plan refuses that flow and to use draft: true. The second is right on Free. The 14 Sep batch used neither — it scheduled live. Whichever autonomy model Dave picks (memo 4), the drafting skill's "never take a post out of draft" line is already untrue of the queue and should be rewritten to describe the batch-approval loop.
2. Metricool tier map — what each tier buys us
Prices from metricool.com/pricing (fetched 15 Sep; the page bills in EUR/USD, GBP figures below are approximate at €1 ≈ £0.86 and are not what the invoice will say). Per-brand prices are the smallest brand count on each tier; we are one brand.
Tier
Price
Brands
Posts
Analytics history
What is gated behind it
Free (today)
€0
1
20/month (help centre, 13 Jul 2026 article)
30 days
5 competitor profiles, 5 AI credits/mo, basic reports
Post approval system, team/client roles, Looker Studio, API (Zapier, Make, MCP), custom report templates, full X analytics, 35 AI credits/mo, Advanced Analytics add-on at €30/mo
Custom
"Let's talk"
n
—
—
White label, account manager, custom AI credits
Now the same table from our side. The rule that prunes it is AGENTS.md #18 — Facebook, Instagram, TikTok, YouTube and nothing else — plus #12 (no robot replies to customers) and D-27 (no Metricool tracker on the site, no SmartLinks).
Capability
Tier
Verdict for ProvenBatch
Publishing to FB / IG / TikTok / YouTube
Free
Have it.
20 posts/month cap
Free
The wall. Two Reel/TikTok pairs a week plus Facebook and Shorts is 30–40 items a month. Metricool counts a multi-network post once, so the 14 Sep batch of 15 posts is ~15 of 20 for September. October at the §3 cadence needs ~24–30. This is what forces Starter, and it forces it in October.
30-day analytics history
Free
The second wall. The §1c baseline vanishes by 15 Oct. Either export weekly to the repo (the loop can) or pay Starter.
Best time to post
Free
Have it; TikTok already returns real data.
Scheduling Reels with trending audio (Instagram Audio API, since May 2026) and Trial Reels
Free
Have it. Already used by social-video (audioConfiguration).
Autolists (evergreen recycling, up to 200 posts per list, circular)
Not stated on the pricing page; the help centre's permissions article implies paid
Worth having — the 33-post library plus the FSA-explainer evergreen set could run as a circular list on a low cadence, freeing the weekly loop to do only new material. Verify by trying to create one on Free; if refused, it rides the Starter upgrade.
Flows (comment keyword → auto-reply + DM with a link, Instagram and TikTok; launched 1 Sep 2026, "free for an initial period", later a paid add-on within Starter and Advanced)
Free now; Starter/Advanced later
The most valuable unused feature, and the one that needs a Dave decision. A viewer comments "LABEL" on a Reel and gets the /trial link by DM. It is a customer-facing automated message, so it sits at the edge of #12: the words are Dave's, written once, sent by Metricool, triggered by the viewer. Memo 4 puts it to him. Needs the IG professional account with a linked Page (have) and a TikTok Business account (have).
Inbox (unified comments/DMs)
Available on Free per the help centre; TikTok inbox on paid
Have enough of it. Replying stays human (#12). Meta Business Suite covers FB/IG free anyway.
Competitor analysis (5 → 100 profiles)
Free → Starter
Five is enough: AllergenKit, FoodCore, CompliChef, Label Daze, Paddl. Only IG/FB/YT support competitor tracking.
AI assistant credits (5 → 20 → 35)
all
Worth £0. The copy is written in the repo by Claude/Grok against our copy rules; Metricool's generator would break them.
Reports (PDF/PPT), custom templates, Looker Studio
Starter / Advanced
Worth £0. The scoreboard lives in beta-pipeline.md and the routine writes it.
Post approval system
Advanced
Worth £0 because Dave's ✅ on a weekly chat card is the approval (memo 4). Paying £46/month to move the tick from chat to Metricool's app is the wrong trade.
API / hosted MCP / Zapier
Advanced
Worth £0 while the claude.ai connector works on Free; becomes the Advanced trigger only if Grok's custom-connector test in §7 fails on Free.
Ads manager: create Meta and Google campaigns from Metricool; monitor TikTok Ads
Not tier-gated on the pricing page
Useful as the read side of the ledger (BSCA*, GAEV* metrics come through the MCP). Campaign creation from Metricool requires the Meta Pixel to be configured first — which we do not have and cannot have without the consent decision — so Google campaigns are created in Google Ads by Dave, Meta boosts in Business Suite by Dave, and Metricool reads all three.
Rejected already — a third-party click store is a D-12 sub-processor disclosure. provenbatch.co.uk/trial is the link-in-bio.
Web tracker (be.js)
any
Rejected already (D-27, PECR).
LinkedIn / X / Pinterest / Threads / Bluesky / Twitch / Google Business Profile
Starter unlocks LinkedIn
Worth £0 (#18). Google Business Profile is a listing, not a social channel, but it is still outside the four and is raised as a question in §8 rather than proposed.
Verdict: Free until the October post cap bites (it will), then Starter monthly (~£17, cancel any month), not annual, until trial→paid is known. Advanced only if the Grok connector test fails. Custom never.
3. The organic engine — what Metricool should run for us
The 31 Jul and 30 Aug research already settled the content shape (short-form, chaos-first, product at the reveal, 2 posts a week beats a blitz, TikTok 21–34s, burned captions, first 1.5 s decide). What was not settled is what Metricool itself automates. Four things:
Cadence and slots from measured data, not the calendar file. TikTok's own best-time curve (§1c) says Wed–Fri at 10:00 / 12:00 / 18:00. The loop asks getBestTimeToPostByNetwork every Monday and places the week's pairs on the top three slots; Facebook text posts go out at 19:45 as now (its curve is unmeasurable); YouTube uploads midday Thursday (search inventory, timing matters least).
Two lanes of content, one of them recycled. Lane A is new material: two Reel/TikTok pairs a week from the social-content.md library and the trend intake (≤2 [trend] ideas, already the drafting skill's rule). Lane B is evergreen: the PPDS / "may contain" / HACCP explainers are true for years. An Autolist holds them as a circular list at one slot a week per network, so the feed never goes quiet when Lane A is thin (Cake International week, Christmas). Every Autolist item must carry the current CTA line — the 13 Aug stale-legal-claim lesson applies doubly to a list that republishes itself, so the loop re-reads the list monthly against the source file.
Trending audio at publish time, never baked in (the two-lane rule in media-and-social.md): Instagram via audioConfiguration (read back the resolved title — the resolver fuzzy-matches), TikTok via the Commercial Music Library at direct-post.
One Flows keyword per campaign. "Comment LABEL and I'll send you the link" is the highest- converting CTA pattern on Instagram in 2026 because a comment is cheap and a DM is personal. Metricool's Comment → DM template does: optional public reply (≤300 chars) → first DM asking for a reaction (opens Meta's 24-hour window) → follow-up DM carrying the link with a button. Sends, interactions and CTR are reported per flow. Decision for Dave (memo 4): the DM text is his, fixed, reviewed once; the trigger is the viewer's. If he says no, the CTA stays "link in bio".
Cost of the organic engine per month: Metricool Free/Starter £0–17; generated b-roll clips ~$0.35–2.80 each, two a week ≈ £8–15; narration inside an existing plan; app footage from the mobile harness free. ≈ £10–35/month.(1 Oct 2026, #2988: the b-roll and narration vendors used for these figures are closed; both lines need a replacement tool.) Everything else is robot time.
What organic can plausibly deliver by December, from the §1c base: TikTok grew from 0 to ~3,100 monthly views in five weeks with eleven videos and no follower base. Holding the cadence and adding Flows, a 3–5× view growth by December is the range the July research's "nobody owns the food-safety-educator seat" finding supports, i.e. 10–15k TikTok views/month, 1–2k Instagram, ~1k Facebook page views — and at a 1–2% profile-to-link rate and 5–10% link-to-trial, that is roughly 5–20 started trials a month from organic alone. It is a forecast, not a promise; the loop's Friday numbers replace it with a measurement within three weeks of launch.
4. Measurement — decision 0, because nothing else can be read without it
Today, deliberately (D-27), the marketing site has no analytics at all; the Google Ads tag is shipped dormant (ADS_TAGS_ENABLED = false, #1822); the site can witness only trial_start_click (launch.js, "the account creation itself happens in the app and is not observable from here"); the app records a business with no origin; and beta_requests.heard_via — the self-reported "how did you hear about us" — retires with /beta on 1 October. So on 2 October we will know how many trials started and nothing about where any of them came from.
Three fixes, ordered by cost and by whether they touch PECR:
#
Fix
Cookie/consent?
Cost
What it tells us
M1
Cloudflare Web Analytics on provenbatch.co.uk — a dashboard toggle (D-27: "one toggle away")
No (cookieless, no identifier)
£0, 5 minutes of Dave
Visits, referrers, top paths, so /trial traffic by source (TikTok vs Google vs direct) becomes visible
M2
Carry origin into the signup record.launch.js already forwards utm_source/medium/campaign/term/content and gclid/gbraid/wbraid to app.provenbatch.co.uk (CARRIED_PARAMS). The app should store them on the new business (one JSONB column, e.g. signup_attribution), and a "how did you hear about us" one-liner on the create-account card replaces heard_via. journey_events then joins trial → activation → paid back to a source
No — reading URL parameters the site itself put there and writing them server-side stores nothing on the visitor's device
One small app issue (schema + card + admin RPC); not filed today
Cost per started trial, per channel — the only number the ROI model needs
M3
Google Ads conversion via the dormant gtag + consent banner (#1822's three-things-in-one-pass), or Meta Pixel/CAPI for conversion-optimised Meta campaigns
Yes — a consent banner on a trust product whose privacy policy §12.1 promises none
Design + legal + the #1822 pass
Lets Google and Meta optimise to signups rather than clicks
A cookieless alternative to M3 for Google exists and is enough at our volume: with M2 storing gclid at signup, conversions can be imported into Google Ads offline (Dave uploads a CSV of gclid, conversion time weekly, or the loop prepares it) — no tag, no banner, and the exact-match campaign still learns which keywords produced trials. Meta has no equivalent without a pixel or CAPI, which is one more reason Meta stays at "boost for reach" rather than "campaign for conversions" in this window.
Recommendation: M1 today, M2 as the one small build this research asks for before 1 October, M3 deferred until trial→paid is measured and the consent question is worth its cost. Memo 2.
5. Paid acquisition — reopened, and modelled channel by channel
5a. The economics that bound every channel (unchanged since gtm-channels-research.md §5a)
Input
Value
Source
Blended ARPU
£14–15/month
D-02 ladder, July mix assumption
Churn
4–5%/month
July research
Gross LTV
£175–280
July research
Allowable CAC (3:1)
£30–100; £60 used as the working ceiling
July research
Trial→paid, opt-in (no card) B2B SaaS 2026
median ~9–14%, range 8–22%; a compliance product with an activation gate (D7 label printed) should sit mid-range — 10 / 15 / 20% low / mid / high used below
ChartMogul 2026 study (8.9% median), First Page Sage (18.2% organic), Growth Spree benchmark table
**⇒ Maximum cost per started trial**
CAC × trial→paid = £6 / £9 / £12 at £60 CAC; £10 / £15 / £20 at £100
arithmetic
Two consequences. First, the 14 Sep kill rule of "> £40 per trial after £100" is generous by a factor of two to four against these economics — it is a stop-the-bleeding rule, not a this is working rule; the continue gate has to be ≤ £20 per started trial and the scale gate has to come from measured trial→paid, not from clicks. Second, at £150/month the whole paid budget buys roughly 5–25 started trials a month if a channel performs at its best case, which is the same order as organic (§3). Paid is a supplement and a learning budget in this window, not the engine.
5b. Channel benchmarks, UK, 2026 (fetched 15 Sep — each row's source in §10)
Channel
Cost basis
Best / mid / worst assumed
Click → started trial
Minimums and constraints
Google Search, exact-match on allergen-label / Natasha's Law / PPDS software terms
CPC; UK B2B SaaS £3–9, long-tail niche terms usually below that; no published CPC for our terms
£2 / £4 / £6
High intent: 20 / 15 / 10%
Volume-capped: these queries are low hundreds a month at most in the UK, so £8–10/day may not spend. Created in Google Ads by Dave (not from Metricool without a pixel). Conversion learning needs M2 + offline import
Meta boost of a proven Reel (Instagram/Facebook, UK, "food business / home baker" interests + creative that self-filters)
Cost per link click; boosts run ~2–3× the CPC of an Ads Manager campaign (a controlled test put boost at $3.09 vs $1.04); UK niche B2B CPC £0.80–1.60 in Ads Manager
£0.60 / £1.20 / £2.00
Cold social: 5 / 3 / 2%
£20–50 per boost; no pixel means optimise for link clicks or engagement only; Meta's 2% UK "location fee" since 1 Jul 2026 sits outside Ads Manager metrics
TikTok Spark Ads (boosting our own best organic TikTok, traffic objective, UK)
£20/day ad-group minimum, £50 lifetime ⇒ a meaningful week is ~£140. Learning phase rarely exits under £50/day; treat as a reach/creative test, not a conversion machine. Created in TikTok Ads Manager (Metricool monitors only)
Meta conversion campaign (cold prospecting to signup)
CPL £20–40 niche B2B; the July research's realistic test floor is £300–600/month for 6–8 weeks
—
needs pixel/CAPI
Outside the ceiling and outside M3 — modelled only as an escalation
Against the £6–20 ceiling in §5a: only the best cases clear it, and only Google's mid case comes close. That is the honest reading and it matches the July verdict — paid at this ACV is a learning budget until something (landing conversion, creative, trial→paid) proves better than the benchmark. The gates below are how the loop finds out cheaply.
5d. Gates the loop applies every Friday (per channel, on M2's attributed numbers)
Gate
Rule
Action
Continue
cost per started trial ≤ £20 over the trailing 14 days and ≥ 3 trials
keep the daily cap
Hold
£20–40, or fewer than 3 attributed trials after £60
keep spending, but the card carries a "fix the landing / swap the creative" item; no increase
Kill
> £40 after ≥ £100, or D7 activation of attributed trials < 20%, or > 50% irrelevant search terms after negatives
pause; Dave is asked, not told
Scale (Dave's decision, never the loop's)
trial→paid on the first cohort ≥ 15% and cost per trial ≤ £20
memo proposing a higher ceiling (the £400/month band)
6. ROI model — the three months after launch
Assumptions: Starter Metricool from October (£17 monthly), generated b-roll ~£12/month, ARPU £14, mid case for every channel unless stated, trial→paid 15% at month 3, no churn inside the window (it is too short to matter). Organic trials from §3's range at the low end. Costs are cash; time is not costed. All figures are forecasts to be replaced by the Friday numbers.
6a. The recommended plan inside £150/month
Month
Spend
Forecast started trials
Notes
Oct
Metricool £17 + b-roll £12 + Google ~£80 (capped £8/day, volume-limited) + one £25 boost of the best launch-week Reel = ~£134
organic 5–8 + Google 2–4 + boost 1–2 = 8–14
Launch week filled by the §7 loop; M1 on, M2 shipped, Flows keyword live if Dave says yes
Nov
Metricool £17 + b-roll £12 + Google ~£80 or one TikTok Spark week £140 (the loop recommends which on 1 Nov from October's cost per trial) + one £25 boost after Cake International = ~£134–195
organic 8–12 + paid 3–6 = 11–18
Cake International 6–8 Nov: post-show boost of the show Reel; the £195 variant needs Dave's explicit nod
Seasonal peak for home bakers; first October trials reach their 30-day decision point
Quarter
~£400–460
~32–53 started trials
at 15% trial→paid ⇒ 5–8 paying businesses by early January; at 10% ⇒ 3–5; at 20% ⇒ 6–11
Payback. 5–8 paying at £14 ARPU is £70–112 MRR against ~£430 spent: 4–6 months to recover the quarter's cash, inside the 22-month expected lifetime, so ROI-positive on the July LTV — and that is at mid-case throughout. It also crosses the July research's break-even of 8 Starter or 4 Standard customers for the fixed cost base only at the top of the range. Growth "fast" in this window means tens of trials, not hundreds; the levers that change that are trial→paid (product, not marketing) and organic reach compounding (Flows, TikTok cadence), not the £150.
6b. The gated escalation ladder (Dave funds it or it does not exist)
Rung
Trigger
Spend
Forecast, mid case
Why it is a rung and not the plan
E1 Raise Google cap to £15/day if it is actually spending £8/day at ≤ £20/trial
Continue gate met two Fridays running
+£200/month
+7 trials/month
Exact-match volume is finite; the first sign it is not is unspent budget
E2 Recurring TikTok Spark, £20/day, best organic video each fortnight
October Spark week ≤ £20/trial and D7 ≥ 20%
+£600/month
+20 trials/month
Learning phase and creative fatigue; needs a fresh winner every two weeks
E3 Meta conversion campaign with pixel or CAPI
M3 decided (consent banner) and trial→paid ≥ 15% measured
£300–600/month for 6–8 weeks
+10–20 trials/month at £20–40 CPL → £30–60/trial — fails the ceiling unless trial→paid ≥ 25%
The July finding stands: cold Meta does not clear the economics at this ACV; it only makes sense as retargeting or lookalikes after ~100 customers
E4 Creator seeding: UK cake-business coaches at 1–10k followers, £20–150/post
Flows live and a proven Reel to hand them
£100–300/month
unmodelled — the July research rates it above cold paid on trust
Not a Metricool feature; listed because it competes for the same £
7. Autonomy and orchestration — the weekly loop
7a. What may run without a human, and what may not
Runs itself
Needs Dave's weekly ✅
Stays Dave's, always
Reading analytics, best times, scheduled queue, ad metrics
Scheduling the coming week's posts live
Anything customer-facing that is a reply (#12) — Inbox stays draft-only
Activating a boost or a Spark week (he clicks, or says "go" and a session with the ad-manager login does)
Ad billing, budget ceilings, new ad accounts
Writing the Friday scoreboard into beta-pipeline.md §D by PR
Changing a Flows message
Facebook groups (no automation, ever)
Autolist rotation, Flows sends (once approved), Metricool's own best-time placement
Raising a kill-gate pause
Taking Metricool off Free
7b. The loop, on the pattern the digest already proves
The beta-updates-digest Grok automation (11 Sep 2026) is the proven shape: a scheduled fire reads the repo and the database, delivers one card in Charlie/Grok chat, and a later fire reads Dave's ✅ / ✏️ / ❌ in that chat and fulfils. Automations cannot be triggered by a reply, only by schedule (or email on SuperGrok), which is why fulfilment is a second scheduled fire.
Fire
When (UK)
Reads
Does
Writes
Draft
Monday 07:30
outputs/gtm/social-content.md, the trend sources in the drafting skill, Metricool: last 7 days' analytics, next 14 days' queue, best times; ad metrics (GAEV*, BSCA*); Cloudflare Web Analytics if M1 is on; the app's attributed trials if M2 is shipped
Composes week+1: 2 Reel/TikTok pairs, 1–2 Facebook posts, Shorts where a clip is <60 s, ≤2 [trend] ideas; names one boost candidate (best reach × CTA-bearing post of the last 7 days) and the Spark/Google verdict against §5d; drafts the Friday scoreboard row
One chat card: full copy, slots, media URLs, the numbers, the ads verdict, and the questions (boost yes/no, Spark week yes/no). No buttons (#1026).
Fulfil
Tuesday 07:30 (and Wednesday as a catch-up if Monday's card has no reaction)
Dave's reply in the chat (his account only)
✅ → schedules the batch live via createScheduledPost (autoPublish: true, draft: false, AIGC flags, YouTube type: "short" only under 60 s); ✏️ → redrafts the card; ❌ → drops. Ads: never changes spend — posts the instruction for Dave
The scheduled posts; a chat confirmation line; ops-report.yml dispatch (job outcome), the same as every routine
Score
Friday 07:30 (fold into Ops: weekly review & health or a third fire)
Metricool analytics, ads metrics, M2 trials
Applies §5d gates; opens a docs-only PR updating beta-pipeline.md §D and, monthly, exports the analytics JSON to the repo so Free's 30-day window stops mattering
PR + one-line chat note
Caps, in the jobs.json style: 1 card per draft fire; ≤ 8 posts scheduled per fulfil; ≤ 2 trend ideas; ≤ 1 boost recommendation; £0 of spend changed by the routine, ever. Decision kind social_ref already exists.
7c. Who runs it — the one open technical question
Option
What it needs
Status
G1 — Grok automation with Metricool as a custom MCP connector (Dave's stated preference: "Grokbot can be employed to orchestrate")
grok.com/connectors → New Connector → Custom → Metricool's hosted MCP URL and its auth. xAI: connectors "available to all Grok users", scheduled Automations "free for everyone on grok.com"; the digest proves GitHub and Supabase are wired for Dave
Untested. Metricool says its MCP needs Advanced; the claude.ai connector says otherwise on Free. A ten-minute test by Dave settles it: if it authenticates on Free, G1 is the plan; if it 403s, either pay Advanced (~£46 monthly) or fall to G2
G2 — Grok drafts and scores; the existing Claude Ops: growth routine (Monday 08:00, holds the Metricool connector) fulfils
The Claude routine's prompt rewritten to read Dave's ✅ from the approval card and schedule live — a prompt edit Dave makes in the Routines UI (a session cannot edit a routine it did not create)
Works today with no new purchase. Two robots, one queue: only the fulfiller may write to Metricool, and the card must say which
G3 — all three fires as Claude Routines
Dave creates them in the claude.ai UI (session-created triggers get no connectors)
Contradicts "Dave declined new Claude Routines after the cutover" and #1189's tension noted in claude-to-grok-transfer-research.md §1c. Listed for completeness
Either way: Grok's TUI scheduler_create is not a Routine (7-day, session-scoped — claude-to-grok-transfer-research.md §2) and must not be used for this. Grok Bot, the separate agent app, is not required: xAI's scheduled Automations are the product the digest runs on, and sources conflict on Grok Bot's price ($200/month standalone versus bundled with SuperGrok) — verify at x.ai before assuming either.
7d. Before 1 October, in order
Dave: M1 toggle; the G1 connector test; the Flows yes/no; the tier yes/no (memo 1, 2, 4).
A session: ship M2 (one small app issue, to be filed only if Dave says so — this research files nothing); rewrite the drafting skill to the batch-approval loop; register the routine in jobs.d/; export the §1c baseline JSON to the repo before it ages out.
The loop's first Draft fire: Monday 28 September, filling 1–7 October — launch-week copy switches from "Opens 1 October" to "Start your free 30-day trial" (launch.js flips the site at 00:01; the queue does not flip itself).
8. Products and services we are not using — considered, with a verdict
Product / service
Cost
Verdict
Metricool Starter
~£17 monthly
Yes, when the 20-post cap bites (October). §2
Metricool Advanced
~£46 monthly
Only if the Grok custom-connector test fails on Free
Metricool Advanced Analytics add-on
+€10/30
No — volume too small to need Campaign Dashboards
Metricool Flows
free now, add-on later
Yes if Dave accepts the one fixed DM (memo 4)
Metricool Autolists
tier-dependent
Yes for the evergreen lane; verify on Free first
Metricool SmartLinks
paid
No (D-12 sub-processor)
Metricool web tracker
any
No (D-27)
Looker Studio, Zapier, Make
Advanced
No — the repo is the reporting layer
Google Business Profile listing (Metricool supports it)
£0
Question, not a proposal. It is a listing, not a social channel, and #18 says four channels and nothing else. If Dave wants "ProvenBatch" to show a Knowledge Panel, it is a five-minute claim; otherwise leave it
Cloudflare Web Analytics
£0
Yes, today (M1)
Google Ads offline conversion import
£0
Yes once M2 stores gclid
Meta Business Suite inbox
£0
Already enough for FB/IG replies (human)
YouTube Studio
£0
Comments and Shorts analytics; nothing to buy
Generated b-roll / narration
per clip
Vendors closed 1 Oct 2026 (#2988); a replacement is Dave's decision. The Friday numbers stay the real virality predictor
Resend (transactional, lifecycle)
free to 3,000/mo
Keep; the trial nurture sequence (day 1 / 7 / 21 / 28) is the highest-leverage email item and is product work, not Metricool — the lifecycle mailer already exists
MailerLite / newsletter tooling
free to 250 subs
No — the updates list is beta_requests + the console's Campaigns card
Social listening (Brand24 ~$199, Mention ~$41, Awario $29–74)
monthly
No — deferred post-launch in RESEARCH §16; nothing changed
Creator seeding (£20–150/post)
per post
Escalation E4 once a proven Reel exists
Grok Bot (agent app)
disputed: $200/mo standalone or bundled with SuperGrok
Not required; Automations are the product in use
9. Recommendation and the 90-day calendar
Do, in this order:
Measurement first (memo 2): M1 today; M2 before 1 October; M3 deferred.
Metricool Starter monthly in October (memo 1), Free until then; nothing above Starter unless the Grok connector needs it.
The paid mix inside £150/month (memo 3): Google exact-match at an £8/day cap from 1 October, one £25 boost of the best launch-week Reel in week two, a single TikTok Spark week in November only if Google under-spends, every channel on the §5d gates. This supersedes the 14 Sep Google-only decision by adding the boost and the Spark test; it keeps Meta campaigns off.
The Grok-orchestrated weekly loop (memo 4): first Draft fire Monday 28 September; Flows keyword if Dave accepts the fixed DM; Autolist for the evergreen lane.
Week of
Milestone
15 Sep
Memos to Dave; M1 toggle; G1 connector test; queue runs to 27 Sep as scheduled
22 Sep
M2 built if approved; drafting skill rewritten; baseline JSON exported; loop registered
28 Sep
First Draft fire fills 1–7 Oct; Metricool → Starter if the cap is hit
1 Oct
Launch. Google exact-match live at £8/day; Flows keyword live; /trial CTA flips at 00:01
5–12 Oct
First Friday score; £25 boost of the best launch Reel
19–26 Oct
First trials reach D7 — activation gate readable
2 Nov
October review: Google continue/hold/kill; Spark week decision; Nov plan
6–8 Nov
Cake International — show Reel; post-show boost
30 Nov
First October trials reach their 30-day decision: first trial→paid reading
7 Dec
Scale/escalation memo to Dave if the §5d scale gate is met; otherwise hold at £150
4 Jan 2027
Quarter review against §6a; formal cohort (D-26) window opens 23 Nov meanwhile
10. What this research did not do, and follow-ups
Did not test the Grok custom connector against Metricool's MCP — Dave's grok.com account is needed. It is the single fact that decides G1 versus G2 and Free versus Advanced.
Did not read Google Keyword Planner (dashboard-only, as the July research already noted); the exact-match volume assumption is "low hundreds a month" and could be lower.
Did not verify Autolist availability on Free by creating one — a write on the live brand.
Did not price Grok Bot from x.ai directly — two 2026 sources disagree; the loop does not need it.
Did not fetchdocs.x.ai/docs/guides/automations (404); the Automations facts come from xAI's connectors pages and three secondary write-ups of the 16 Jul 2026 launch.
GBP prices for Metricool are conversions, not invoice figures; the pricing page bills EUR/USD.
Follow-ups that would harden this most: the first three Friday scoreboards (they replace §3 and §6's forecasts); an M2-attributed cost per trial by channel; a trial→paid reading on 30 Nov.
Sources (all fetched 15 Sep 2026 unless dated)
Live Metricool brand: getBrandSettings, getScheduledPosts, getAnalyticsAvailableMetrics, getAnalyticsDataByMetrics, getBestTimeToPostByNetwork via the claude.ai Metricool connector.
Metricool pricing: metricool.com/pricing · metricool.com/premium-vs-free-metricool-plans (13 Jul 2026) · help.metricool.com "Fair Use Policy for Social Media Scheduling" (600 posts/brand/month) · help.metricool.com "Your guide to the Advanced Analytics add-on" · metricool.com/mcps-for-marketing ("Advanced plan or higher") · help.metricool.com "How to manage Flows", "Flows on Instagram", "Flows on TikTok" · metricool.com/press-release-metricool-flows (1 Sep 2026: free for an initial period, later a paid add-on within Starter and Advanced) · help.metricool.com "Ads Campaigns Management" and "Create ads campaigns from Metricool" (Google and Meta creation; Meta Pixel required; TikTok monitored) · help.metricool.com "Schedule content from an Autolist" (200 posts, circular) · help.metricool.com "Best time to post" · metricool.com "How to add audio to your Instagram Reels" (Instagram Audio API, 18 May 2026) · releasebot.io Metricool notes to 11 Sep 2026 (Reporting section, Trial Reels with trending audio).
xAI / Grok: docs.x.ai/grok/connectors ("Connectors are available to all Grok users"; New Connector → Custom → MCP server URL) · x.ai/news/grok-connectors (6 May 2026: GitHub, Google Workspace, Bring Your Own MCP) · MindStudio, Basenor and AI/TLDR write-ups of Grok Automations (16 Jul 2026: schedules free for everyone; email triggers on SuperGrok) · poster.ly Grok Bot guide (18 Aug 2026: remote MCP servers only, up to 50 routines) · aibuilderclub.com Grok Bot pricing (30 Aug 2026) and vellum.ai breakdown (20 Aug 2026) — these two disagree on price.
Trial→paid: ChartMogul 2026 study via kirro.io and growthspreeofficial.com (opt-in median 8.9%, range 8–22%); First Page Sage via userpilot.com (18.2%).
Meta: gtm-channels-research.md §5b (UK CPC £0.80–1.60, CPL £20–40, 2% UK location fee from 1 Jul 2026) · admetrics.io boost-vs-Ads-Manager test ($3.09 vs $1.04 per click) · enrichlabs.ai 2026 benchmarks (leads CPC $1.92, CPL ~$27.66) · growthspreeofficial.com B2B SaaS Meta benchmarks 2026.
Do not add competitor write-ups under outputs/docs/. Edit the sheet and the Drive documents instead.
Regulation
Last updated 19 Jul 2026
4. Regulation (UK)
Natasha's Law / PPDS — in force since Oct 2021
Food packed on site before being ordered or selected must carry the food name and a full ingredients list with the 14 regulated allergens emphasised. This is the app's whole reason to exist and is already implemented.
Ingredient list — how nested/compound ingredients must be declared (researched 19 Jul 2026)
Dave asked the exact question: do we merge sub-ingredients, name only the ingredient, or list them nested? Answer: it's a hybrid, and the full definitive write-up (with worked example, app mapping and gaps) now lives in bakery-manager-labelling-solution-design.md §1A — read that before touching label generation. Short version:
The finished label is one flat list in descending weight order (Reg 1169/2011 Annex VII Part A).
Bought compound ingredients (chocolate chips, digestives, baking spread…) are named with their sub-ingredients in brackets — Name (a, b, c) — Annex VII Part E. (Dave's "nested" option.)
Home-made components / sub-recipes (our own caramel, biscuit base) are flattened to their raw ingredients and merged into the list by weight — you don't name them. (Dave's "merge" option.)
"Name only the ingredient" is wrong except for genuinely single-component items — it hides allergens.
All 14 allergens must be listed AND emphasised, always — even inside a compound under the 2% breakdown exemption. We always show the full breakdown (never rely on the exemption).
The app already does the hybrid correctly (Case A via Name (declaration), Case B via expandToLeaves flatten from v0.18.0). Open gaps to backlog: QUID %, net quantity, and "Use by" vs "Best before" for perishable cream/cheesecake items. See design doc §1A.
Precautionary allergen labelling ("may contain")
PAL is voluntary, not a legal requirement, and should only follow a real risk assessment of unavoidable cross-contamination that can't be controlled by segregation and cleaning.
FSA best practice: name the specific allergen — "may contain peanuts", not "may contain nuts".
⚠️ Implication we have not resolved: we auto-propagate a supplier's "may contain" onto our labels. That is a defensible default, but it is not a risk assessment, and it can produce generic "may contain nuts" wording that FSA guidance discourages. Worth a decision.
Watch: FSA is considering standardising PAL, with Codex thresholds (ED05). UK needed a position for CCFL May 2026; Codex adoption was on the agenda for July 2026; any UK change is not expected before late 2026. This is the live regulatory item for us.
Benedict's Law — does NOT apply to us
Corrected 17 Jul 2026. Our labelling design doc said to "monitor Benedict's Law guidance (expected Sept 2026)" as though it affected our labels. It doesn't.
It became law in England in March 2026 via the Children's Wellbeing and Schools Bill, and requires schools to hold an allergy policy, an allergy risk register, spare adrenaline auto-injectors and staff training from 1 Sept 2026. It is education-sector legislation and explicitly does not extend to hospitality or food labelling. Trade press confusion on this point is common. Action: correct the labelling design doc.
Nutrition
Believed exempt for small producers (micro-enterprise exemption), which is why nutrition (#28) is ranked low. Unverified — confirm before building #28.
Found while cleaning footnote markers off pack transcriptions (#256, #257).
Retained EU Regulation 1333/2008 Annex V requires six colours to carry the statement "may have an adverse effect on activity and attention in children": E102 (Tartrazine), E104 (Quinoline Yellow), E110 (Sunset Yellow FCF), E122 (Carmoisine/Azorubine), E124 (Ponceau 4R), E129 (Allura Red AC). Named after the 2007 University of Southampton study that prompted the rule.
This is not an allergen rule, so nothing in our allergen detection catches it — a separate labelling obligation riding on the ingredient list. Source: Reg 1333/2008 Annex V as retained in UK law; FSA guidance on food additives.
Evidence from our own data: two of Dave's colourings (sourcings 53, 55) contain E122, and the packs mark it with an asterisk pointing at the printed warning. Our transcription kept the pointer and dropped the target — which is how the gap surfaced at all.
Worth checking before building #257 whether the same pattern applies to other families: aspartame E951 ("contains a source of phenylalanine"), polyols ("excessive consumption may produce laxative effects"), high-caffeine drinks, liquorice. If so, model it generally as "additive → required statement" rather than special-casing colours.
Regulation
Last updated 30 Jul 2026
8e-1. It DOES generalise — the full family of mandatory statements (#257, researched 30 Jul 2026)
Answering the question §8e left open. Yes: model it generally. But the finding that matters is not the list — it is that the app can only ever decide a third of it, and the reason is structural rather than a gap we can close by trying harder.
Two regulations, not one
The colours rule is the odd one out. It lives in Reg 1333/2008 Annex V (food additives). Every other mandatory statement lives in Reg 1169/2011 Annex III (FIC), "Foods whose labelling must include one or more additional particulars". Both are retained UK law. A model built only around "additives" will not naturally reach Annex III, because two of its entries are not about additives at all — so the concept to build is "finished food → required statement", keyed by whatever triggers it, not "additive → statement".
The full set, and what triggers each
#
Trigger
Required statement
Where it must go
Colours
any of E102, E104, E110, E122, E124, E129 present
"may have an adverse effect on activity and attention in children"
with the colour in the ingredients list
Sweeteners
any authorised sweetener present
"with sweetener(s)"
with the name of the food
Sweeteners + sugar
both added sugar(s) and sweetener(s)
"with sugar(s) and sweetener(s)"
with the name of the food
Aspartame
E951 / aspartame-acesulfame salt present
"contains aspartame (a source of phenylalanine)" if declared by E number; "contains a source of phenylalanine" if declared by name
after the ingredients list
Polyols
more than 10% added polyols (E420, E421, E953, E965, E966, E967, E968)
"excessive consumption may produce laxative effects"
after the ingredients list
Liquorice (low)
glycyrrhizinic acid / ammonium salt / Glycyrrhiza glabra at ≥100 mg/kg (confectionery) or ≥10 mg/l (beverages)
"contains liquorice" — not needed if "liquorice" already appears in the ingredients list or the name
🔴 The finding that should shape #257: three kinds of trigger, and we can only compute one
Presence-only — colours, aspartame, sweeteners, protective atmosphere. Detectable from the declaration text we already store. The app can decide these.
Concentration in the finished food — liquorice (mg/kg, mg/l), caffeine (mg/l). Needs the quantity of the substance in the finished product, which we do not hold and cannot derive: a declaration says what is in a pack, never how much, and the recipe knows pack weights rather than the additive's concentration inside them. The app can never decide these.
Percentage of the finished food — polyols >10%. Theoretically derivable from recipe quantities if we knew the polyol content of each pack. We do not. Cannot decide.
So the honest design is detect presence, then ask — for two of the three classes the app's role is to raise the question and record the human's answer, not to reach a verdict. That is the same shape as the existing allergen rule (the machine may detect, only a human may accept), which is convenient, but it needs saying out loud because a naive build would try to compute thresholds it has no inputs for and would then be confidently wrong in the silent direction — the #204 class.
A second finding: the label renderer has no concept of placement zones
Three distinct placements appear above — with the name, immediately after the ingredients list, and in the same field of vision as the name. Our label output has no notion of any of them; it has an ingredients block and an allergen emphasis. Whatever #257 builds needs somewhere to put a statement, and "after the ingredients list" is the only one the current layout can express. Worth deciding deliberately rather than discovering at render time.
A third: widening beyond bakers (D-24) pulled two new entries into scope
Neither would ever have mattered to a baker, and both are ordinary for businesses we now say we serve: date of freezing on frozen meat and meat preparations (a butcher), and "packaged in a protective atmosphere" (a deli or farm shop that MAP-packs). This is a small, concrete example of the positioning decision having a compliance consequence rather than just a copy consequence.
Relevance ranking for our actual users
Real today: the six colours — evidenced in Dave's own catalogue (sourcings 53, 55). Plausible soon: sweeteners, aspartame and polyols, via sugar-free icings, "no added sugar" ranges and protein bars. Narrow but real: liquorice (confectioners), date of freezing (butchers), protective atmosphere (delis). Effectively never: phytosterols.
⚠️ Sourcing caveat — read before writing any of this wording onto a label
⚠️ CORRECTED 25 Aug 2026 — the fetching half of this caveat is FALSIFIED. Re-tested from a cloud session, all four sites answer:
Source
Then
Now (25 Aug 2026)
legislation.gov.uk (plain page)
403
HTTP 200
legislation.gov.uk (data.xht?view=snippet)
403
HTTP 200 — verbatim statutory text
fsai.ie
403
HTTP 200
businesscompanion.info
403
HTTP 200
eur-lex.europa.eu
403
HTTP 202 — a processing response, not a refusal
The original entry is kept below because a falsified reason invites a bad "correction" later, and because the network genuinely did behave that way when it was written — the environment's egress policy was widened around 9 Aug 2026, which is the likeliest cause. What changes as a result: this is no longer off-keyboard work. A session can now pull the retained text itself, so the character- for-character confirmation below should be done in-session and not handed to Dave. ⚠️ The substantive half of the caveat stands unchanged and is the part that matters — see the paragraph after the original.
Original entry (superseded, kept as the record):The primary legislative text could not be fetched from a cloud session.legislation.gov.uk, eur-lex.europa.eu, fsai.ie and businesscompanion.info all answer HTTP 403 to an automated fetch — the agent proxy was verified healthy at the time (recentRelayFailures: []), so this is the sites refusing bots, not our network.
The part that still holds, whatever the network does. The wordings and thresholds above come from secondary sources, cross-checked between several. That is good enough to design against and not good enough to print: statutory wording is exact, and a near-miss is a non-compliant label. Before #257 ships, each statement must be confirmed character-for-character against the retained text on legislation.gov.uk.
14. HACCP / Safer Food, Better Business (SFBB) — the daily due-diligence norm for #36 (researched 9–10 Aug 2026)
Read directly for #36's Food safety build (outputs/docs/overnight-batch-issues-35-36-plan.md §C.10) — the app's first primary-source pass on this topic; RESEARCH.md had no HACCP section before this entry. England & Wales only — CookSafe (Scotland) and Safe Catering (Northern Ireland), the two regional equivalents CLAUDE.md already names, were not independently checked this session; do not assume they match SFBB's structure without reading them. ✅ DISCHARGED 6 Sep 2026 — both packs read in full for #1349: see §14c. This warning was right: neither pack has SFBB's daily opening/closing walk, CookSafe verifies weekly against its House Rules and reviews six-monthly, and Safe Catering uses a proprietor's hygiene inspection with a three-item 4-weekly review inside it. This paragraph stays as written because it is what stopped a session relabelling the SFBB seventeen as "CookSafe"; read §14c before touching anything nation-specific.
Primary sources read, verbatim where quoted below:
The GOV.UK publication landing page, gov.uk/government/publications/safer-food-better-business- sfbb/safer-food-better-business-sfbb (published 5 Jun 2025) — a shell page, exactly the food.gov.uk → GOV.UK trap CLAUDE.md and outputs/gtm/social-content.md §2a already flag: it names the pack's five sections and applicability but carries no prescriptive detail itself, linking out to the real documents as separate PDFs.
The SFBB introduction PDF (sfbb-introduction_2.pdf, 9 pages) — how the pack and diary work, registration duty, the Food Hygiene Rating scale, and the "Working with food?" hygiene factsheet.
The SFBB caterer's diary PDF (sfbb-diary-07-diary-and-4-weekly-review-fixed_0.pdf, 7 pages) — the actual daily record template and the 4-weekly review sheet.
Both PDFs failed to extract via WebFetch (returned corrupted/binary text) and via pdfplumber (this environment's cryptography package has a broken native _rust/cffi binding, which pdfplumber/pypdf both pull in transitively and crash on import even for an unencrypted PDF). pymupdf (import fitz) worked — a pure-C-extension reader with no cryptography dependency — and is the one to reach for first if this recurs. Worth recording: this is an environment-level Python packaging defect, not a content problem, and cost real time to diagnose.
What SFBB actually prescribes, and where our #36 seed matches or diverges from it:
The diary is coarser than our schema assumed. Each day has exactly two attestations — "Opening checks" and "Closing checks" — each a single tick plus a name and signature under one sentence: "Our safe methods were followed and effectively supervised today." It is not a per-item log of individual readings (no "fridge temp: 4°C" line in the diary itself). The itemised detail — what each "safe method" actually checks, including any numeric limits for chilling/cooking/hot-holding — lives in the separate sector-specific "safe methods" sheets (e.g. the full caterer's pack), which this session did not fetch; that stays unverified, not guessed at. #36's seed splits "Fridge temperature" and "Freezer temperature" out as their own named checks rather than folding them into one bundled "Opening checks" tick — a reasonable product decision (a due-diligence app can usefully separate what a paper diary bundles), but worth knowing it's our structure, not a literal transcription of SFBB's. ⚠️ PARTLY CORRECTED 26 Aug 2026 — see §14b, which reads the pack this session skipped. The two-attestation finding is right about the diary page, but the caterer's pack prints the seventeen itemised opening and closing checks those two ticks attest to (pp. 66 and 82), so itemising them is a literal transcription of SFBB after all, not our own structure. Only the temperature split (fridge/freezer as named checks rather than 'prove it' readings) is ours. Read the way it stood, this entry invites a future session to "simplify" the itemised list back into one bundled tick on the grounds that SFBB bundles it — SFBB does not.
⚠️ CORRECTED 26 Aug 2026 (#730/#645 corpus review) — the reasoning below was falsified, the decision itself was not re-opened. The original entry said "no numeric temperature limit appears in either PDF read this session… none confirmed here against a primary FSA safe-method sheet" — true only of the two documents actually read (the intro and the diary), not of SFBB as a whole. The caterer's "safe methods" pack — sfbb-caterers-pack-fixed_0_3.pdf, live at assets.publishing.service.gov.uk (re-fetched and text-extracted with pymupdf, the #36 pack's own PDF-reading fix, since pdfplumber/WebFetch both still choke on this file) — does state primary numbers, verbatim: "Frozen food should be kept at -18°C or below"; food "can be put back in the fridge and kept at 8°C or below"; hot-held food must stay "at a safe temperature of 63°C or above"; and cooking gives explicit time/temperature combinations — "80°C for at least 6 seconds, 75°C for at least 30 seconds, 70°C for at least [longer]…" — not a single ≥75°C figure, which the commonly-repeated shorthand flattens. So the widely-cited 8°C / ‑18°C / 63°C figures ARE genuine primary FSA text, just from a document this session hadn't read yet. Decision §C.9 (ship the seeded checks with no default number) is UNCHANGED by this — #730 explicitly separates "the decision still looks right" from "the recorded reason for it was wrong," and a wrong reason sitting next to a right decision is exactly what invites a future session to "helpfully" reverse the decision on false grounds. If a default number is ever added, the cooking one specifically needs the time/temperature table, not a flat "≥75°C".
"Extra checks" (SFBB's term) are irregular, non-daily checks — maintenance, deep-cleaning a freezer — logged once at the end of the week, not daily. Distinct from our four_weekly frequency, which maps more directly to the 4-weekly review below.
The 4-weekly review is a walk-round + a fixed yes/no checklist, not a diary entry. It asks, among other things: "Have probes been calibrated in the last 4 weeks and results recorded?" — this is the primary-source anchor for #36's seeded "Probe thermometer calibration" check, though SFBB frames it as one item on a 4-weekly review, not a standalone repeating check. #36 seeded it at weekly frequency (a reasonable product simplification — calibrate more often than the legal minimum is conservative, not wrong) but if this is ever revisited, four_weekly is closer to the primary source's own cadence.
Corrections are explicitly a narrative, not a structured record: "If things go wrong, write down what happened and what you did in your diary." SFBB's own due-diligence model is closer to a free-text incident note than #36's structured amend-with-supersession — our design is stricter (a real correction trail, enforced at the RLS layer) than the paper form it's replacing, which is a reasonable direction to diverge in, not a mismatch to fix.
Who signs: "the person responsible for running the business," daily — matches #36's recorded_by capturing whichever team member is signed in, not a fixed single named operator.
Registration and Food Hygiene Rating are out of scope for #36 (they're a one-off business registration duty and an EHO-assigned outcome, not a repeating check) but are useful context: using the pack correctly is explicitly stated as improving the Food Hygiene Rating outcome, which is the commercial argument for #36 existing at all.
How the cadence is anchored: rolling from when you start, NOT a fixed calendar or week numbers (researched 16 Aug 2026, for the #558-UAT question "how is a weekly/4-weekly check's due day determined?"). The FSA guidance is explicit that the 4-weekly review is done "at the end of every 4 weeks" — you look back over the past 4 weeks' diary and complete it even if no problem was found. There is no FSA-published national calendar, week-numbering scheme, or fixed reference date: the 4 weeks run rolling from whenever a business starts, and 4 weeks is 28 days, deliberately not a calendar month (a month would drift between 28–31 days and lose the fixed-length review window). The daily diary is filled in each day; "weekly" extra checks are logged by the end of the week. Implication for our scheduling model: a business-set week-start day defining the week a weekly check must be completed within, plus 4-week blocks measured off that same grid, is faithful to SFBB — whereas ISO week-number parity (which starts weeks on Monday regardless of the business's own week-start) would diverge from it, so it was not adopted.
14a. HACCP plan authoring — the GOV.UK/MyHACCP primary-source pass for #440 (researched 12 Aug 2026)
Scoping pass for #440 (#36 phase 3); design decisions in outputs/docs/issue-440-haccp-plan.md. Extends §14 — same England & Wales framing, same "no numbers found, so ship none" outcome.
Read directly (verbatim quotes in the #440 doc):
food.gov.uk/business-guidance/hazard-analysis-and-critical-control-point-haccp now 301-redirects to gov.uk/food-safety-management-systems — the food.gov.uk→GOV.UK migration §14 flagged has since swallowed the HACCP business-guidance page itself. The GOV.UK page states "If you run a food business, you need to have a food safety management system" and lists seven requirements to meet HACCP principles (identify hazards; identify CCPs; set limits for the CCPs; monitor the CCPs; put checks in place to make sure your system is working; put things right if there is a problem with a CCP; keep records). No numeric temperature or critical limit appears anywhere on it — consistent with §14's finding for the SFBB diary, and again the reason the app ships structure without numbers.
Known only via search extracts from FSA's own site — NOT a page read:
MyHACCP (myhaccp.food.gov.uk) — the FSA's free HACCP-plan authoring tool, aimed at UK food manufacturing businesses with 50 or fewer employees; guides a business through 8 preparatory stages (A prerequisites, B management commitment, C scope of study, D select the team, E describe the product, F identify intended use, G flow diagram, H confirm flow diagram) and the 7 principles (1.1 identify hazards / 1.2 hazard analysis / 1.3 control measures; 2 determine CCPs; 3 critical limits; 4 monitoring; 5 corrective action plan; 6 validation, verification and review; 7 documentation), producing a downloadable PDF of the study. ⚠️ The site answers HTTP 403 to every non-browser fetch from this environment (WebFetch and curl with a browser UA alike — a WAF, not an outage; the FSA's acss.food.gov.uk copy of the authorised-officer guidance PDF is a 404 now too). The structure above comes from FSA-site search extracts, so treat exact stage/principle wording as unverified until someone reads the site in a real browser. #440's design deliberately doesn't depend on the wording — it ships no MyHACCP content, only a structure compatible with it.
The scope fact that shapes #440: SFBB is the FSA-approved food safety management system for caterers and retailers — a business running SFBB properly needs no bespoke HACCP plan; MyHACCP is the FSA's route for small manufacturers. Our PPDS users straddle the line, so plan authoring is an optional capability and the app must never imply the SFBB-style daily-checks half (#36 phases 1–2) is insufficient on its own.
Still unread, still not guessed at (unchanged from §14): the SFBB sector safe-methods packs; CookSafe (Scotland); Safe Catering (NI).
MyHACCP (FSA) — 403 to non-browser fetches; structure via FSA-site search extracts, 12 Aug 2026
Regulation
Last updated 26 Aug 2026
14b. The SFBB opening/closing checks — the verbatim primary source behind the #534 seed (researched 26 Aug 2026)
⚠️ This section was CITED before it was WRITTEN.index.html's SFBB_OUTLINE_CHECKS comment, safety_reviews' comment, goaudits-fsa-checks-ux-research.md and road-to-launch-overnight-batch- plan.md all pointed at "RESEARCH.md §14b" as the primary-source authority for what the seeded checks say — and no §14b existed. The seed's contents were correct (verified below, item for item), but for ten days the evidence for the app's only verbatim regulatory transcription was unciteable from the repo. Written here to close that, prompted by a Whisk & Whimsy observation (below) that sent a session looking for it.
Source read in full: the FSA Safer Food, Better Business for caterers pack, sfbb-caterers-pack-fixed_0_3.pdf (100 pages), fetched from assets.publishing.service.gov.uk and text-extracted with pymupdf — still the only reader that works on these files in this environment (WebFetch and pdfplumber both fail; §14 records why). The opening/closing lists appear twice, identically: as the safe method sheet (p.66) and again in the diary's own instructions (p.82). Quoted verbatim.
OPENING CHECKS — "You should do these checks at the beginning of the day. You can also add your own checks to the list." (nine items)
"Your fridges, chilled display equipment and freezers are working properly."
"Your other equipment (e.g. oven) is working properly."
"Staff are fit for work and wearing clean work clothes."
"Food preparation areas are clean and disinfected (work surfaces, equipment, utensils, etc.)."
"All areas are free from evidence of pest activity."
"There are plenty of handwashing and cleaning materials (soap, paper towels, sanitiser, etc.)."
"Hot running water is available at all sinks and hand wash basins."
"Probe thermometer is working and probe wipes are available."
"Allergen information is accurate for all items on sale."
CLOSING CHECKS — "You should do these checks at the end of the day." (eight items)
"All food is covered, labelled and put in the fridge/freezer (where appropriate)."
"Food on its Use By date has been thrown away."
"Dirty cleaning equipment has been cleaned or thrown away."
"Waste has been removed and new bags put into the bins."
"Food preparation areas are clean and disinfected (work surfaces, equipment, utensils etc.)."
"All washing up has been finished."
"Floors are swept and clean."
"'Prove it' checks have been recorded."
So SFBB is itemised, and the paper diary is not — both are true, and the distinction matters. §14 recorded that the diary page carries just two attestations ("Opening checks" / "Closing checks", one tick each). That is the signature line. The seventeen items above are what those two ticks attest to, printed on the facing instructions page and again in the safe method. A product that offers one bundled "hygiene check" is therefore reproducing the signature and discarding the checklist — the wrong half. Our per-item model is the faithful one, and #534's decision to itemise was right for a reason stronger than the "reasonable product decision" §14 hedged it as.
Our seed vs. the source — 17 of 17 present, three deliberate trims.SFBB_OUTLINE_CHECKS carries every opening and closing item in the source order, plus three temperature checks the FSA lists as 'prove it' readings rather than walk-round ticks (fridge morning/evening, freezer daily) and the weekly probe calibration (§14 explains why weekly not four_weekly). Trimmed for row width, with no loss of scope: "chilled display equipment" (folded into Fridges and freezers working), "probe wipes" (folded into Probe thermometer working), and "'Prove it'" → Records completed. The §C.9 posture holds throughout — names only, no numeric target, no hazard, no citation in the UI.
Whisk & Whimsy, 26 Aug 2026 — the legacy-seed gap this surfaced. W&W reported "several hygiene sub-checks, not one" during opening checks. They were right, and the cause is data, not code: their list is the pre-#534 six-check auto-seed (10 Aug), which bundled the whole walk into a single "Opening hygiene check" and "Closing hygiene check". #560's repair backfilled their food_safety_seeded_at on 16 Aug; #534's itemised outline shipped two days later, on 18 Aug, gated on an empty check list. Net effect: every business seeded before v0.210.0 is permanently locked out of the itemised outline — the "Start from a typical SFBB outline" button renders only in the empty state, and foodSafetySeedOutline() returns early on a non-empty list. Not a regression in #534, and correct for its own idempotency purpose; it simply has no top-up path. Production blast radius on 26 Aug 2026 is one business (W&W) — it is the only tenant with any checks at all — but it is Dave's own pilot, i.e. the one account whose EHO pack would be looked at first. Filed as its own issue; the repair is additive (offer the missing items, never silently insert them).
14c. CookSafe (Scotland) and Safe Catering (Northern Ireland) — the primary-source pass §14 said was owed (researched 6 Sep 2026)
Read for #1349, which makes the UK nation a stored fact (business_settings.nation) and seeds the food-safety module from the right pack. §14 flagged this gap on 9 Aug 2026 in as many words — "England & Wales only — CookSafe (Scotland) and Safe Catering (Northern Ireland) ... were not independently checked this session; do not assume they match SFBB's structure without reading them" — and that warning turned out to be the important sentence in the whole section. They do not match SFBB's structure. Neither pack has SFBB's daily opening/closing walk at all.
Sources read in full, both fetched as PDFs and text-extracted with pymupdf (still the only reader that works on these in this environment — WebFetch returns the GOV.UK shell page and pdfplumber crashes on import; §14 and §14b record why, and it cost nothing this time because they did):
Safe Catering — your guide to making food safely, FSA in Northern Ireland, Issue 7, April 2019, 130pp — assets.publishing.service.gov.uk/media/69b3f32ffdbfc4d58fc8cf2e/safe-catering_0_0.pdf, plus its recording forms (safe-catering-recording-forms_0_0.pdf), SC5 and SC8 as separate PDFs.
⚠️ food.gov.uk/business-guidance/safe-catering 301s to GOV.UK — the same food.gov.uk → GOV.UK trap §14 flagged. The GOV.UK page is the one that carries the "Applies to" statement.
The three packs are siblings, and the FSA says so itself
Safe Catering's own foreword, verbatim: "Other similar guidance materials have also been developed by the Food Standards Agency i.e. Safer Food Better Business developed in England and CookSafe in Scotland." That is the authority for #1349's nation → pack mapping; it is not our inference.
GOV.UK on Safe Catering: "Food safety management guides for caterers and retailers in Northern Ireland", and explicitly "Applies to Northern Ireland".
🔴 Safe Catering is a JOINT FSA-NI / Food Safety Authority of Ireland product, which matters for #826 rather than for #1349: "This joint initiative with the Food Safety Authority of Ireland is intended to assist with consistency of application of food hygiene legislation right across the island of Ireland." That sentence is from the FSA-NI publication. The FSAI does publish its own Safe Catering Pack. The FSAI publishes its own Safe Catering pack with the same SC-numbered forms.(Struck 15 Sep 2026, #1879. Un-struck in part 16 Sep 2026, #1524, having been read form by form from the FSAI's own sheets — the resolution is below and the full read is §29d.)
🔴 RESOLVED, THREE WAYS — and "same SC-numbered forms" was two-thirds right (#1524, 16 Sep 2026).
The SC numbering IS used on the Irish side. #1879's strike was against the wrong artefact. It read the FSAI's Record Books web page, which titles the sheets "Recording Form n". The sheets themselves print their SC number on the artwork — SC4, SC5, SC6, SC7 and SC9 were all extracted directly from the PDFs — and the FSAI's own asset filenames are SC2-2012-refrigeration.pdf, SC5-2012-inspection.pdf, SC6-2012-training.pdf and SC7-2012-fitness.pdf. A page title is not the form.
The nine slots match FSA-NI's SC1–SC9 one for one, in order, by purpose. Two titles differ: FSAI 2 is Refrigeration (NI: Fridge/Cold room/Display Chill Temperature Records) and FSAI 9 is Transport and Delivery (NI: Customer Delivery Record).
So the Northern Ireland profile #1349 built was indeed most of an Irish profile — the shape and the numbering transferred, the rows did not.
CookSafe (Scotland) — structure
"'CookSafe' is designed to assist catering businesses understand and implement a HACCP based system." Five sections: Introduction, Flow Diagram, HACCP Charts, House Rules, Records.
There is no daily opening/closing walk. The periodic verification is the Weekly Record: "The following ongoing checks should be carried out by the Manager or Proprietor during each working week and should be carried out by all businesses using 'CookSafe'." It asks "Have the House Rules been followed? YES / NO / N/A" against ten headings — Training, Personal Hygiene, Cleaning, Cross Contamination Prevention, Pest Control, Waste Control, Maintenance, Stock Control, Temperature Control, and Records ("Have all necessary Temperature Checks been recorded using the correct recording form/s?") — then "If the answer to any of the above questions is 'NO' then enter the corrective action details in the table below" under HOUSE RULES DEVIATIONS OBSERVED / CORRECTIVE ACTIONS TAKEN. It is signed once: "Manager/Proprietor's Signature ... Date".
Records section, its own numbered list: "1. Delivery Record 2. Cold Food Record 3. Hot Temperature Record 4. Hot Holding Record 5. Off Site Temperature Record 6. All-In-One Record (used as an alternative to records 1 - 5) 7. Weekly Record". The All-In-One is "To be completed daily". Cold units: "Temperature checks (Recommended twice daily)" with AM and PM columns; freezers get "Function checks (Recommended once daily)" — a function check, not a temperature reading, which is why #1349's CookSafe seed carries "Freezer working" as a tick and not a temperature.
Probe thermometer: MONTHLY, its own sheet — "The readings in iced water should be -1°C to +1°C", "The reading in boiling water should be between 99°C and 101°C", plus "The electronic display unit should be checked at least once per year."
Review: SIX-MONTHLY, not four-weekly."It is essential that your HACCP based procedures are kept up to date. A review of your system must be carried out on a regular basis, ideally every six months or if any of the circumstances covered in the table below arise." The table is a list of triggers (new dish, new equipment or supplier, premises changes, House Rules changes, "A Local Authority inspection where deficiencies were noted", new hazard information, cleaning chemical changes, staff changes, customer complaint), not a fixed questionnaire.
Safe Catering (Northern Ireland) — structure
"developed principally for catering businesses, but it may also be used by retailers who have a catering function within the business." HACCP's seven steps, then a Declaration of Completion of your Safe Catering Plan signed and dated, with future review dates recorded on it.
No daily opening/closing walk here either. The premises walk is SC5 – Hygiene Inspection Checklist: "Simple checks of the premises which should be carried out by the Proprietor or Manager regularly", where "regularly" is the business's own choice — the form's footer is "Tick frequency checks carried out by proprietor or manager: Weekly / Fortnightly / Monthly". Its sections are Hygiene of Food Rooms & Equipment, Food Storage, Food Handling Practices, Personal Hygiene, Pest Control, Waste Control, Checks and Record Keeping, and then a Review (4 weekly) block of exactly three questions: "Any new suppliers and approved list updated?", "Any new menu items and steps in Safe Catering updated?", "Any new food handling methods or equipment and steps in Safe Catering updated?" Signed with Name / Position / Signed / Date.
Daily records are SC1–SC4, foldable into SC8 – All-In-One Record: "This form may be completed daily and used as an alternative to the individual records". SC2's footnote sets the floor: "It is recommended that fridge temperatures are checked at least once per day", on a form with AM and PM columns. SC4 is "(For Food To Be Held Hot For More Than 2 Hours)".
Safe Catering DOES state primary numbers, unlike CookSafe (which leaves critical limits blank for the business to write in): "Chilled food: max. 8°C; Hot Food: minimum 63°C" (SC1), "Core temperature above 75˚C" (SC3), "Keep hot food above 63˚C" (SC4). Recorded here for completeness; §C.9 is unchanged and #1349 ships no number in any seeded check name, exactly as §14/§730 settled for SFBB.
The daily sign-off: SFBB has one, the other two do not
SFBB's diary carries a daily attestation — "Our safe methods were followed and effectively supervised today" (§14). Neither CookSafe nor Safe Catering has a daily declaration sentence at all. CookSafe's signature is on the weekly record; Safe Catering's are per-form ("Manager/Supervisor check on ___ Initials") plus the one-off plan Declaration. This is why #1349's FOOD_SAFETY_PACKS carries an attributed flag and why the app now says "the end-of-day declaration" rather than "the SFBB daily declaration" to a Scottish or Northern Irish business: the app's own daily sign-off is a good product decision either way, but attributing its wording to a pack that does not contain it would be a fabricated citation.
🔴 CORRECTED IN PLACE, 16 Sep 2026 (#1524) — that last sentence describes only HALF of what shipped, and the half it leaves out is the citation. Found in browser UAT (outputs/mobile-harness/uat1524.js), not by reading the code. attributed governs the caption under the sign-off card; the quoted sentence above it is SAFETY_SIGNOFF_STATEMENT, a plain constant (part-1.js), and it is SFBB's line rendered in quotation marks to every business in every nation. So a Scottish business is shown, today, on main:
"Our safe methods were followed and effectively supervised today." The end-of-day declaration. Signing records your name and the time…
The caption stopped naming SFBB; the quotation marks did not stop quoting it. This is the exact #22 shape the attributed flag was introduced to prevent, in the line it was introduced for — a fix that looks complete because the thing it renamed is the thing anyone would check.
Not fixed by #1524, deliberately, and this is the reasoning rather than an omission. It affects GB-SCT and GB-NIR on main right now; changing it changes what an existing, non-IE business sees, which epic #827's dark-groundwork posture forbids this issue from doing (#1897: "OFF = BYTE-IDENTICAL TO TODAY"). It is #1349's own defect and wants its own issue, its own UAT and its own copy decision — an unattributed wording that works for all four nations, or a per-pack sentence with attributed: false printing none. Raised on #1524 rather than filed silently.
Fixed 18 Sep 2026 (#1930), option (b). §14c does not supply a genuine all-pack sentence — neither CookSafe nor Safe Catering has a daily declaration at all — so inventing one would be the same fabricated citation the attributed flag was introduced to stop. The quoted sentence is now FOOD_SAFETY_PACKS.sfbb.signoff, printed only where attributed: true, and omitted for CookSafe, Safe Catering and FSAI. Both the sign-off card (.fssignoff-stmt) and the EHO pack meta go through foodSafetySignoffStatement(), so they cannot disagree. The caption already behaved this way; the quote now matches it.
⚠️ The label does NOT differ by nation — verified, not assumed
The issue asked for this to be checked rather than believed, because the Windsor Framework keeps Northern Ireland under EU food-information law. Nothing on a PPDS label differs across the four nations. PPDS took effect 1 October 2021 across the whole UK; the 14 allergens are identical; emphasis within the ingredients list is required identically. What differs is only which agency publishes the guidance: GOV.UK's PPDS labelling guidance states "Applies to England, Northern Ireland and Wales" and points to Food Standards Scotland for the Scottish publication of the same requirement. The corpus already recorded this (ppds-natashas-law.md: "No difference in the substance of the requirement was found between the four nations, only in which agency publishes the guidance") and this pass confirms it against the regulators' current pages.
Consequence for the build, and it is the important one:business_settings.nation is read by the food-safety seed and by ProvenBot, and by nothing on the labelling path. A future session must not "extend" nation-awareness into the label engine — there is nothing there to vary, and a nation-split label would be a regression, not a feature. outputs/tests/uknations1349.js pins this.
What the app cannot yet express, recorded so it is not rediscovered
CookSafe's six-monthly review.SAFETY_FREQUENCIES stops at four_weekly; there is no monthly or six-monthly frequency, and adding one is a CHECK-constraint migration on two databases plus a drift_info() change. #1349 therefore seeds no review check for Scotland rather than inventing a cadence the pack does not have.
CookSafe's monthly probe check is seeded at four_weekly (28 days) — the nearest available and a slightly shorter interval than the pack asks for, which is the safe direction to round in.
Safe Catering's weekly/fortnightly/monthly choice on SC5 is seeded at weekly, the most frequent option it offers and the only one of the three the app can express. A business that inspects monthly edits the row.
The guided 4-weekly review (SFBB_REVIEW_ITEMS) is still SFBB's twelve questions for everyone. #1349 makes only its name nation-neutral. Rewording it per nation needs its own issue: CookSafe's equivalent is six-monthly and trigger-based, and Safe Catering's is the three-item block above.
GOODSIN_REF_NOTE still cites the SFBB pack's delivery temperatures to every business. CookSafe states no equivalent figure (the business writes its own critical limit) and Safe Catering's SC1 states chilled max 8°C and hot min 63°C but no frozen-delivery figure, so there is no verified per-nation replacement to swap in. Left as it is, deliberately, and filed separately. ⚠️ Half of that is no longer true for IRELAND (#1524, 16 Sep 2026). The FSAI's SC9 note states all three: "Chilled food: 0°C to 5°C; Frozen food: less than or equal to -18°C; Hot food: greater than or equal to 63°C", and SC2 repeats the first two. So an Irish replacement is verified and is stricter than the Northern Irish one on chilled food (0–5 °C, not max 8 °C). #1524 ships no number (§C.9) and does not touch GOODSIN_REF_NOTE, which is a UK-wide string outside its scope — but a future issue swapping that note per nation now has the Irish figures and should not re-derive them, and must not assume Ireland and Northern Ireland share them.
This is the Ireland research pass epic #827 and children #1519–#1525 cite. A first pass was written 9 Sep 2026 as RESEARCH.md §27 / §27a–§27g on origin/claude/intl-expansion-epic-split-0pvtix. That branch was deleted unmerged. On main, §27 is GitHub Actions cost (13 Sep) and §28 is used twice (card payments; Metricool). This section is the next free number, rebuilt 15 Sep 2026 (#1879) from the issue bodies and re-read against primary sources the same day. Until this section landed, the approved design record (outputs/docs/ireland-jurisdiction-profile-design-record.md, #1519) was the source for every Irish regulatory claim. The design record is the spec; this section is the research record.
Every regulatory and tax claim below has a primary source and a date. Claims from the 9 Sep issue bodies that did not survive the 15 Sep read are struck in place at the end of this section (and in §11 / §14c where they had already been copied).
How FSAI Guidance Note 28 was read (15 Sep 2026)
The issue that filed this recreate (#1879) inherited a 9 Sep warning that GN28's body was scanned images. That warning is wrong for the current file.
gn-28-food-allergen-declaration-for-non-prepacked-foods-in-ireland-rev-2.pdf (fetched 15 Sep 2026, 283 KB, PDF 1.6)
How read
Downloaded from fsai.ie; text extracted with pypdf (PdfReader.extract_text over all 23 pages). The body is real, selectable text, not scanned images. OCR was not used and was not needed. The same file is what the FSAI publication page links as the current revision.
Not relied on
Revision 1; any Irish-language edition; model memory of the text.
Flag
Page 1 of the body (the contents page) still carries the words "Draft / Not for publication". It is nonetheless what fsai.ie serves as Rev 2. Treated as published guidance because it is the regulator's current file; #1519 §10.2 still gates the label on a re-check of that publication page before #1520 lands.
Sources read this pass (15 Sep 2026)
Source
How read
Date on the source
S.I. No. 489 of 2014 — Health (Provision of Food Allergen Information to Consumers in respect of Non-Prepacked Food) Regulations 2014
Irish Statute Book HTML, irishstatutebook.ie/eli/2014/si/489/made/en/print (browser UA; the site refuses a non-browser agent). Regulations 1–18 and the Schedule.
Made 23 October 2014; in operation 13 December 2014
S.I. No. 656 of 2024 — the Amendment Regulations
Same route, eli/2024/si/656/made/en/print, all ten regulations and the explanatory note.
Made 25 Nov 2024 (Iris Oifigiúil 26 Nov 2024 per #1519; explanatory note read 15 Sep)
PDF fetched; text extracted with pypdf; §2.2 read in full.
"Document last reviewed May 2025"
Revenue — VAT thresholds
revenue.ie/.../vat-thresholds.aspx.
Published 6 May 2026
Revenue — Current VAT rates
revenue.ie/.../current-vat-rates.aspx.
Published 1 January 2026
FALCPA / 21 U.S.C. § 343(w)
Cornell LII HTML of 21 U.S.C. § 343, subsection (w). FDA Food Allergies page fetched the same day for the consumer-facing "Contains" example.
Statute as amended by the FASTER Act (sesame, 1 Jan 2023); FDA page retrieved 15 Sep 2026
Reg (EU) 1169/2011 is cited through GN28 Rev 2 (which reproduces Art. 21 and the Commission Notice of 13.7.2017 in §12), not re-read from EUR-Lex in this pass.
29a. Ireland treats PPDS food as non-prepacked — there is no Natasha's Law
The claim #811, #825, #827 and every child asserted from 28 Aug 2026 holds, and has now been read from the instrument rather than from the FSAI's paraphrase.
S.I. 489/2014 Reg. 3, verbatim: "These Regulations apply to all food which is not prepacked food and is offered for sale or supply, including supply free of charge, to the final consumer or to a mass caterer, including— (a) food packed at a food business operator's premises at the consumer's request, and (b) food packed for direct sale or supply."
That is the UK's PPDS population, named in the Irish instrument as non-prepacked. The FSAI allergen page says the same in its own words: non-prepacked food includes "foods packed on the premises for direct sale to the consumer or mass caterer, e.g., lasagne made in a café kitchen and sold packaged from a fridge in the café" (fetched 15 Sep 2026).
What the operator must do — Reg. 4: written particulars of any allergen, at the point of presentation, sale or supply. Not a full ingredient list. Reg. 5(1) sets the manner (accessible before sale; English, or Irish and English; conspicuous; legible; no confusion as to which food). Reg. 6 exempts a separate declaration where the name of the food clearly refers to the allergen.
The 14 allergens are the same as the UK's — Reg. 2 points at Annex II of Reg (EU) 1169/2011 as amended by Delegated Regulation (EU) 78/2014; GN28 Rev 2 Annex I reproduces them.
The word "Contains" is guidance, not statute. GN28 Rev 2 §9: "In the absence of a list of ingredients, food allergen information must be provided in written format by including the word "Contains…" followed by the particular allergen(s) (e.g. Contains wheat, barely, soya and egg)." [sic — "barely" is the FSAI's typo]. The FSAI allergen page: "It is not acceptable to say 'Our food contains…'. You must identify the exact food, e.g., 'spaghetti bolognaise - contains milk, celery, wheat'." The statute's own words are Reg. 5(1)(e): no possibility of confusion as to which food.
The sentence that shapes the Irish label — GN28 Rev 2 §9:"Non-prepacked food usually does not require a list of ingredients, but where one is provided, any food allergens used in the production or preparation of that food must be emphasised in accordance with the requirements of the FIC Regulation and may not be repeated elsewhere outside of the ingredients list." So an Irish label is one of two shapes, never both: a "Contains:" statement with no ingredient list, or an ingredient list with allergens emphasised and no separate allergen line. The UK PPDS label prints both. That combination is what Irish guidance rules out.
QUID, date line, net quantity are Art. 9 / 22 / 24–25 particulars of prepacked food. They are not required for this declaration. (What the app prints anyway is a #1519 decision, not a research finding.)
What we now hold (26 Sep 2026, #1926). The Irish ProvenBot corpus at outputs/provenbot/knowledge/ie/labelling-allergens.md (plus the home-business and Safe Catering files) states this rule positively, summarised from the FSAI pages re-fetched the same day and from GN28 Rev 2. §29a is no longer only a record of what the UK corpus lacks. Licence and shape: §29h.
What S.I. 656/2024 actually changed: nothing about the declaration. Read from the Statute Book 15 Sep 2026, it inserts the Official Controls Regulation definition, substitutes sampling / second-expert-opinion / sample-division / compliance-notice / service-of-notice provisions, and its own explanatory note says it aligns "compliance notices and sampling" with other 2024 official-controls amendments. Regs 3–6 of S.I. 489/2014 are untouched and have been in operation since 13 December 2014. The FSAI page cites both instruments because together they are the "Regulations 2014 and 2024" (S.I. 656/2024 Reg. 1(2)). Citing "489 as updated by 656" is correct as a citation and wrong as a reason the declaration moved.
29b. US "Contains:" crossover — FALCPA (for #1519 / #1529)
The Irish non-prepacked declaration is structurally the same derivation as the US "Contains:" statement: named food + the word "Contains" + the allergen list. They differ in which allergen set applies, whether species must be named, and where the statement sits relative to an ingredient list.
21 U.S.C. § 343(w)(1) (FALCPA, as amended) — a food is misbranded if it is, or contains an ingredient that bears or contains, a major food allergen, unless either:
(A)"the word "Contains", followed by the name of the food source from which the major food allergen is derived, is printed immediately after or is adjacent to the list of ingredients (in a type size no smaller than the type size used in the list of ingredients)"; or
(B) the source name follows the ingredient in parentheses in the list (with the usual exceptions where the ingredient's name already uses the source name).
The FDA's own consumer page (fetched 15 Sep 2026) gives the example "Contains wheat, milk, and soy." The FASTER Act added sesame as the ninth major food allergen from 1 January 2023. Tree nut, fish and Crustacean shellfish must be named by type/species (§ 343(w)(2)) — the same species_named idea GN28 §12 requires for cereals and nuts in Ireland.
The shared derivation is allergenDeclaration()'s statement mode (named food + lead + allergens). Ireland prints the statement instead of a list; the US prints it after the list (or in the list). Whichever of #1520 / #1529 lands first builds that mode for both; the second must not fork it. The allergen set (eu14 vs the US nine) is #1528's data, not this derivation.
FALCPA's requirements apply to packaged foods offered for sale; they "do not apply to foods that are placed in a wrapper or container … following a customer's order at the point of purchase" (FDA Food Allergies page, 15 Sep 2026). That is the US analogue of "not this declaration" — recorded so #1529 does not treat FALCPA as covering every counter sale.
29c. Half the profile is already built — prepacked is EU FIC
LABEL_FORMATS = ["ppds", "prepacked"] with resolveLabelFormat(): the app already has two formats, and the second is EU FIC, which the UK retained essentially intact. An Irish business selling genuinely prepacked food is already in the right regime; that half is checked for Brexit-era divergence (FBO address must be EU-established — an Irish address is; language — English satisfies; allergen emphasis — identical), not rebuilt. The one divergence GN28 Rev 2 §9 forces on the Irish FIC row is dropping the separate Allergens: line (declaration: none), because an emphasised list's allergens "may not be repeated elsewhere". #1519 decides the rows; this section records the research that made the reuse cheap.
29d. FSAI Safe Catering — and the joint-product claim from the Irish side
The FSAI publishes the Safe Catering Pack as "a practical, easy-to-use, food safety management system", "designed for caterers but it may also be used by other food businesses", "ideal for businesses that have not yet developed their own food safety management system" (Safe Catering Pack page, fetched 15 Sep 2026). It contains a plan the operator completes, hygiene requirements, a Declaration of Completion and Review, and recording forms. A 31 July 2026 news item added sections on food donations, water supply and food safety culture.
The Record Books page (8 July 2022) lists nine forms: (1) Food Delivery, (2) Refrigeration, (3) Cooking/Cooling/Reheating, (4) Hot Hold Display, (5) Hygiene Inspection Checklist, (6) Hygiene Training, (7) Fitness to Work Assessment, (8) All-in-one daily record, (9) (the page continues; the takeaway-compliance page names SC1–SC9 including Transport and Delivery as SC9). Individual forms can be downloaded free; the pack can be purchased.
HACCP has been a legal requirement for all Irish food businesses since 1998 (FSAI HACCP-based Procedures page, fetched 15 Sep 2026). The Safe Catering Pack is the FSAI's route to one — the same relationship SFBB has to the UK obligation.
29d-1. The form-by-form read — done (#1524, 16 Sep 2026)
⚠️ The paragraph this replaces said the SC-numbering claim "is not confirmed from the Irish side" and left it to #1524. It is now confirmed, partly. Kept as a heading of its own rather than edited away, because the shape of the answer — two-thirds right — is the useful part.
How it was read. All nine recording forms were downloaded from the FSAI's own Record Books page on 16 Sep 2026 (fsai.ie/publications/safe-catering-pack-record-books → fsai.ie/getmedia/…) and their text extracted directly from the PDFs. FSA-NI's safe-catering-recording-forms_0_0.pdf (Issue 7, April 2019) was fetched the same day and diffed against them, so the comparison is sheet against sheet rather than sheet against §14c's prose. Both sets are linked at the end of §14c.
⚠️ What could NOT be read, said plainly. FSAI forms 1, 3, 4 and 8 (Food Delivery; Cooking/Cooling/Reheating; Hot Hold Display; All-in-one) are served as outlined artwork — their titles, SC numbers and a couple of printed example rows extract, their column headers and footnotes do not. Everything asserted about those four below comes from the Record Books page's own one-line description of each, and nothing else. Forms 2, 5, 6, 7 and 9 extracted in full.FSAI_OUTLINE_CHECKS therefore gives the rows derived from forms 1, 3 and 4 no hint at all — a hint is the source's own examples in plain English (#777), and writing one for a form nobody has read is the fabricated-citation move §14c refused for the daily declaration.
Fridge/Cold room/Display Chill Temperature Records
title differs
3
Cooking, Cooling, Reheating
Cooking/Cooling/Reheating Records
✅
4
Hot Hold Display
Hot Hold Display Records
✅
5
Hygiene Inspection Checklist
Hygiene Inspection Checklist
✅
6
Hygiene Training
Hygiene Training Records
✅
7
Fitness to Work Assessment
Fitness to Work Assessment Form
✅
8
All in one Daily Record
All-in-one Record
✅
9
Transport and Delivery
Customer Delivery Record
title differs
The SC numbering is the FSAI's own, not FSA-NI's read back: SC5 and SC4 are printed on the artwork, SC9 - Transport and Delivery Records and SC6 - Hygiene Training Record extract as text, and four of the FSAI's asset filenames carry the SC number. §14c is corrected accordingly.
SC5 is where the two packs actually part company. Its seven section headings — Hygiene of Food Rooms & Equipment, Food Storage, Food Handling Practices, Personal Hygiene, Pest Control, Waste Control, Checks and Record Keeping — and its three-item "Review (4 weekly)" block are word for word the same in both. Its rows are not. FSA-NI Issue 7 rebuilt Food Handling Practices around a "clean area" concept the FSAI sheet has no trace of, and adds:
"Is outer packaging removed from ready-to-eat food before being placed into a clean area?"
"Are suitable BS EN approved cleaning chemicals available…" (FSAI: no BS EN)
"Are separate cleaning cloths used in clean areas? If they are re-used are they laundered in a boil wash?" (FSAI: "Are cleaning cloths suitable for use and regularly cleaned and disinfected and used properly?")
"Are wash hand basins clean with hot water…" (FSAI: warm water)
"Is a separate probe thermometer used for ready-to-eat [food]…" (FSAI: "Are probe thermometers correctly used and cleaned/disinfected before and after use?")
Both carry "Are staff aware of food allergy hazards?", and neither pack's seed carries it as a row — recorded so that is a decision rather than an oversight. SC5's own header actor differs too: FSAI "carried out by the Manager or Supervisor regularly", FSA-NI "by the Proprietor or Manager". Both footers tick Weekly / Fortnightly / Monthly, so #1524 seeds weekly for the same reason #1349 did — the most frequent option offered and the only one of the three the app can express.
🔴 THE THREE DIFFERENCES THAT REACH THE SEED (the rest are recorded, not built):
Freezer temperature is expected, not optional. FSAI SC2: "This record book should be used for recording temperatures of fridges, freezers etc … It is recommended that fridge/freezer temperatures are checked at least once per day." FSA-NI SC2: "Some businesses may wish to record freezer temperatures." So the Irish outline carries a Freezer temperature row and the Northern Irish one does not. Both sheets carry AM and PM columns, so both seed a morning and an evening fridge reading.
SC5's rows differ (above), so every hint in FSAI_OUTLINE_CHECKS is written from the FSAI's own rows. The seven item names coincide with Safe Catering's because the seven sections genuinely are the same — forcing them apart would have been invention.
The temperature figures differ, and Ireland's are stricter. FSAI SC9 and SC2: "Chilled food: 0°C to 5°C; Frozen food: less than or equal to -18°C; Hot food: greater than or equal to 63°C." FSA-NI SC1: "Chilled food: max. 8°C; Hot Food: minimum 63°C", and no frozen figure at all (§14c). Nothing is seeded from this — §C.9 — but it closes half of §14c's GOODSIN_REF_NOTE gap.
Fixed 18 Sep 2026 (#1930). Browser UAT on #1524 found the daily sign-off still quoted SFBB's declaration verbatim to every nation, Ireland included once enabled — only the caption beneath it was attributed-aware. That was #1349's defect, live for Scotland and Northern Ireland, and #1524 left it (dark-groundwork: changing it would have changed what an existing non-IE business sees). The quoted sentence is now absent for the FSAI pack (and CookSafe / Safe Catering), matching attributed: false. Corrected in place in §14c under The daily sign-off.
Verdict for the build: the FSAI outline is built from the FSAI pack, and it differs from #1349's Northern Ireland branch by exactly one row name, every hint, the regulator and the site. FOOD_SAFETY_PACKS.fsai carries attributed: false for the same reason CookSafe and Safe Catering do: the FSAI pack's declaration is the one-off Declaration of Completion and Review signed when the plan is written, not a daily attestation, so the app says "the end-of-day declaration" plainly.
Not seeded, and why: forms 6 (Hygiene Training) and 7 (Fitness to Work) are personnel records rather than periodic monitoring, and form 9 (Transport and Delivery) covers deliveries to customers — #1349 seeded none of their Northern Irish equivalents either, and parity is deliberate. Form 8 folds 1–4 into one sheet, which is a recording convenience, not a different set of checks.
Sources for this pass (all fetched 16 Sep 2026):
FSAI — Safe Catering Pack — €100, Republic of Ireland shipping only; five new sections (Allergens, Acrylamide, Food Safety Culture, Food Donations, Water Supply) downloadable free
The FSAI also publishes HACCP Catering, Safe Food To Go, and Starting a Food Business at Home — the last is squarely ProvenBatch's customer. Which of those the seed uses is #1519 / #1524, not this section.
29e. Irish tax figures, and the €10,000 threshold that is not ours
Concrete figures, so #1523 can suppress the UK ones honestly rather than inventing an Irish Accounts tab (out of scope):
UK (what the app shows today)
Ireland
Source, dated
Tax year
6 April – 5 April
calendar year
standing Revenue practice (Form 11 is a calendar-year return)
Self-assessment
SA103
Form 11; pay-and-file 31 October (ROS extension 18 Nov 2026 for the 2025 return — as recorded 15 Sep 2026 in #1519 from Revenue's own pages)
#1519 §8 S6; not re-fetched as a TDM in this pass
VAT registration threshold (domestic SME)
£90,000
€85,000 goods (where ≥90% of turnover is goods) · €42,500 services
Revenue What are the VAT thresholds?, published 6 May 2026
Standard VAT rate
20%
23% (also 13.5 / 9 / 4.8 / 0)
Revenue Current VAT rates, published 1 January 2026 — standard 23% effective 1 January 2026
Do not read Revenue's €10,000 threshold as ours. TBE manual §2.2 (last reviewed May 2025), verbatim: the €10,000 B2C place-of-supply exception is available only where "the supplier is established, has a permanent address or is residing in only one Member State". And: "The €10,000 threshold does not apply to supplies of TBE services made by a supplier who is not established in the EU." The same sentence is on Revenue's Electronically supplied services page (published 6 July 2026) and on the VAT-thresholds page (6 May 2026), which adds that the €10,000 intra-Community distance-sales / TBE threshold "only applies where the supplier is established and has their permanent address, or usually resides, in only one Member State. Otherwise, the supplier must register for Irish VAT in respect of such supplies."
ProvenBatch is a UK sole trader, not established in the EU. Irish VAT applies from the first euro of a B2C digital sale to a customer in Ireland. That is why #825 §4 made the overseas billing route a gate, and why D-30 picked Stripe Managed Payments rather than non-Union OSS (decisions/1536-overseas-billing-route.md). The €10,000 figure sitting next to the €85,000 / €42,500 domestic thresholds is the intra-EU micro-business one.
The 3.5% Managed Payments add-on (on top of processing + Billing, on the VAT-inclusive amount) is a billing-route finding, recorded in 1536-overseas-billing-route.md §1 / §4 — it was never this allergen section. It is not restated here.
29f. Registration — what a refusal (or a corpus) has to get right
You have to register your food business before you start operating, even if operating from home. FSAI Register Your Food Business, fetched 15 Sep 2026. Starting a Business: Key Facts: a person proposing to operate a food business "is legally obliged to register with or seek approval from a competent authority prior to commencement of trade, and failure to do so is an offence." The EU obligation underneath is Article 6(2) of Regulation (EC) No 852/2004 (cited on the FSAI hygiene-of-foodstuffs page).
For restaurants, delis, retailers, mobile stalls and home businesses, that is a notification to the HSE through the local Environmental Health office — not a UK local-authority registration. (DAFM / SFPA for other establishment types.) FSAI Register and Competent Authorities pages, fetched 15 Sep 2026.
Starting a Food Business at Home is the FSAI document aimed at ProvenBatch's customer (fetched 15 Sep 2026): home does not exempt the operator from Reg. 852/2004, HACCP, training or traceability.
⚠️ Struck, 15 Sep 2026:There is no charge. The business goes on the food business register with a unique reference number and receives an Acknowledgement of Notification of Food Business Establishment letter. The FSAI register page confirms where (HSE Environmental Health) and when (before operating). It does not state that registration is free, and it does not mention an acknowledgement letter. #1525's registration copy says only what the FSAI page says (#1519 §10.4).
ProvenBot's UK corpus is wrong about Ireland in a specific way: the PPDS / Natasha's Law file describes a full-ingredient-list rule Ireland does not have (§29a). A UK answer to an Irish user asking "what goes on my label?" does not merely cite the wrong regulator — it describes obligations that do not exist. #1525 owned the refusal. #1926 replaced that refusal with an Irish corpus (outputs/provenbot/knowledge/ie/); the registration file still says only what this page says, including the 15 Sep strike of fees and acknowledgement letters. The PPDS / Natasha bans stay. Licence: §29h.
29g. English alone satisfies the language rule — no i18n work
S.I. 489/2014 Reg. 5(1)(b): the written particulars must be "at least in the English or in the Irish language and in the English language". GN28 Rev 2 §9: "at least in the English language, or Irish and English languages, although other languages may also be used in addition to, but not instead of English."English alone satisfies it. This epic opens no Irish-language workstream.
29h. FSAI reuse terms (26 Sep 2026) — the licence that shapes the Irish corpus
#1926's first question: the UK ProvenBot corpus is reusable because Crown copyright material is published under the Open Government Licence v3.0. The FSAI is not a Crown body and OGL does not apply to it. The finding, from pages fetched 26 Sep 2026:
Website pages — reuse is permitted, with conditions.FSAI — Open Data and the Re-Use of Public Sector information (fetched 26 Sep 2026) states that the FSAI complies with the regulations on the Re-use of Public Sector Information and encourages the re-use of the information it produces. Verbatim: "All of the information featured on our website is the copyright of the FSAI unless otherwise indicated. You may re-use the information on this website free of charge in any format." Re-use includes copying, issuing copies to the public, publishing, broadcasting and translating. Conditions, as that page lists them: acknowledge the source and FSAI copyright; reproduce the information accurately; do not use it in a misleading way; do not use it for the principal purpose of advertising or promoting a particular product or service; do not use it for illegal, immoral, fraudulent or dishonest purposes. The page cites the Open Data Directive (EU) 2019/1024, transposed by S.I. 376/2021.
What that means for the corpus. Website pages are summarised and cited under the PSI terms, with source and copyright acknowledged on every Irish file. GN28 is summarised and cited rather than reprinted, because its own imprint asks for a Communications Unit application before reproduction. That is the conservative shape the terms actually permit. OGL v3 is not claimed anywhere on the Irish files.
Register page re-check (26 Sep 2026).FSAI — Register Your Food Business still says "You have to register your food business before you start operating, even if operating from home." Registration is through the local environmental health office of the HSE, DAFM, or the Sea-Fisheries Protection Authority. It still does not state a fee, a wait, a reference number, or an acknowledgement letter. The 15 Sep strike in §29f stands.
Labelling page re-check (26 Sep 2026).FSAI — Labelling / Allergens still lists food packed on the premises for direct sale (lasagne-in-the-café-fridge) as non-prepacked, still requires a written named-food declaration, and still says "It is not acceptable to say 'Our food contains…'". QR codes alone still do not comply with S.I. No. 489 of 2014.
Claims struck or corrected this pass (15 Sep 2026)
Corrected in place, not deleted — a struck wrong belief stops us rediscovering it.
FSAI Guidance Note 28's body is scanned images and could not be text-extracted.Rev 2 (25 June 2025) is real extractable text. OCR not used. The 9 Sep / #1879 "scanned images" premise was wrong for the file fsai.ie currently serves. (Rev 1 was not re-fetched; it is not the current revision.)
Guidance Note 28 Revision 1 is the current guidance.Revision 2, 25 June 2025.
S.I. 656/2024 updated the allergen-declaration rules. It updated sampling, second expert opinion, compliance notices and service of notices. Regs 3–6 of S.I. 489/2014 are unchanged since 13 December 2014. The FSAI's "updated by" citation remains a correct citation.
The FSAI pack uses the same SC-numbered forms as FSA-NI Safe Catering.Not verified from the FSAI's own pages (§29d; struck in §14c).
Registration is free and an acknowledgement letter is issued.Not on the FSAI pages this pass read (§29f).
Revenue's €10,000 threshold is a threshold we can use.It is the intra-EU micro-business place-of-supply exception. A third-country supplier has no threshold (§29e).
What survived, and is now sourced: PPDS-as-non-prepacked; no full ingredient list; "Contains" tied to the named food; English alone; QUID / date / net quantity not mandatory for this declaration; 14 allergens identical; Safe Catering is the FSAI HACCP pack; Irish VAT 23% from the first euro for a non-EU B2C TBE supply; the US "Contains:" statement is the same derivation with a different slot layout.
The spec that consumes this section is outputs/docs/ireland-jurisdiction-profile-design-record.md (#1519, approved 15 Sep 2026). Do not re-derive the label rows here.
Regulation
Last updated 15 Aug 2026
GoAudits + Labl.it + FSA — UX & layout research for our Food safety checks feature
Started 15 Aug 2026 (GoAudits + FSA). Expanded 15 Aug 2026 with Labl.it and a dedicated Layout & UX section, at Dave's request.
Research commissioned by Dave: dig into GoAudits (goaudits.com — a mature digital inspection/audit app that does FSA-style food-safety checks) and Labl.it (labl.it — a labelling competitor) for interface/workflow/layout lessons, and extract from the official FSA what a food-safety check record actually requires. All strands feed the Food safety feature we already ship (#36 phases 1/2/4 + #440 HACCP authoring).
This doc is the product/UX synthesis. The durable factual record (all source URLs, verbatim FSA quotes) lives in RESEARCH.md §14b (FSA primary-source closure), §15 (GoAudits benchmark) and §15a (Labl.it + layout/UX); this doc does not repeat every citation.
Posture reminder (unchanged): phases 1–2 shipped "names only, no numbers" and #440 ships "structure, never content." Nothing below proposes seeding a user's checks with a pre-filled temperature limit or auto-ticking anything (DESIGN-LANGUAGE "nothing may ever tick itself"). What the FSA research does change is that we now hold primary-source, verbatim numbers — so any reference figure we ever surface (help copy, a "typical outline", an EHO explainer) can be sourced to the FSA caterers pack rather than a food-safety blog.
1. The three products in one line each
GoAudits — mobile-first inspection/audit platform (the SafetyCulture/iAuditor competitor): checklist → paged mobile inspection → auto PDF report → corrective action → dashboard trend. Enterprise-leaning, ~$10–30/user/mo, no free tier. No purpose-built allergen or SFBB workflow — food safety is generic checklist content. Its interaction & layout design is the thing to learn from.
Labl.it — a hardware-first subscription prep/date-label printer (6.75" touchscreen device
web admin), £28–34/mo inc. device, aimed at commercial kitchens (Hilton, Domino's, Turtle Bay). Concierge onboarding ("Menu Analyser" — their team pre-loads your products). No due-diligence / diary / FSA / HACCP product at all (only a passive print-log and a coming TempLog® probe). Its lessons are onboarding, the author-vs-reuse split, passive logging and permission tiers, not screen layout (much of its UI is on a physical device we can't inspect).
ProvenBatch (us) — Food safety = safety_checks (name / frequency / tick|temperature + targets), a Today strip of what's due, four views (Today's checks / History / Corrective actions / HACCP plans), append-only entries with a correction trail, corrective actions (#182), and an EHO pack export (#441). No photos, no scoring, no signatures, no push.
The competitive picture that matters: the whole due-diligence/FSA-checks category is white space against Labl.it, and GoAudits treats it only generically. We are already the most FSA-specific of the three. So the job is not to copy a category-leader — it's to borrow GoAudits' proven field ergonomics and Labl.it's onboarding/reuse thinking, on top of a feature that already has the right compliance bones.
2. GoAudits — the interaction ideas worth stealing (ranked)
2.1 Failure reveals the fix, inline — "answer No → the corrective-action form appears"
GoAudits' single best idea. A corrective action ("Action Plan") is not a separate menu: answering a question No makes the fix button appear right there, bound to the failing item. The form is exactly four fields — follow-up notes · assignee (with a preconfigured default) · priority (Low/Med/High) · due date — then Save. It surfaces in the report and the Actions tab.
Our state: we have Corrective actions as a first-class view (#182 "detect, don't decide"), and an out-of-range temperature entry already offers one. But raising it is a hop to a separate tab.
Lesson: surface the "put it right?" prompt inline at the failed/out-of-range check in the Today strip, pre-linked to it. Small change on machinery we already own. Highest value-to-effort.
2.2 "Fail gates evidence" — the highest-leverage compliance-quality trick
A response option can be configured so answering it forces a comment and/or photo before you can proceed ("auditors cannot move on… unless they insert a comment or add a photo"). The app blocks progress rather than letting an inspector skip evidence on a fail.
Lesson: when a check is recorded as failed/out-of-range, require the correction note (and, later, a photo) before the entry is accepted. This is fully compatible with our append-only model — it just makes the "what did you do?" narrative the FSA diary demands non-optional on a fail, which is exactly the record an EHO wants.
2.3 Photo-as-evidence, attached to the item, flowing into the report
Every question can carry a photo (capture/upload, with an on-device annotation canvas — draw, save, and the annotation persists into the PDF). Reviewers cite instant photo-embedded reports as the top strength.
Our state: tick / temperature only; no image capture anywhere; the EHO pack has no photo to print.
Lesson: optional photo on a check entry (and on a corrective-action resolution) materially strengthens the EHO pack — the evidence an inspector values under "confidence in management" (§6.3). Real storage cost; scope properly.
2.4 The report is the compliance evidence, one tap
Branded PDF at completion — cover page (logo + business + checklist + date), score/grade, per-section results, annotated photos inline, corrective-actions list, a summary field + signature pad, timestamp & geo — auto-emailed to recipients.
Our state: the EHO pack (#441) is the right idea already; we print, we don't auto-distribute.
Lesson: we're aligned on concept. Incremental adds: photos in the pack, a signed daily sign-off (§7.2), and share/email (lower priority).
Also aligned/small: dedicated Temperature field type (we have it — correct), voice-to-text comments (a genuine field-ergonomics win for a floury kitchen), N/A that drops out of any denominator (only if we ever score — we shouldn't, §5).
3. Labl.it — a different competitor, and what it teaches
Labl.it turned out not to be a self-serve PPDS design app — it's a subscription thermal-label device for kitchen prep/date labels, with allergens attached to finished products (not derived from ingredients) and set up for the customer by Labl.it's team. It has no recipe/ingredient graph, no real PPDS retail labels yet ("coming soon"), and no due-diligence/FSA/HACCP product. That makes most of our core — and all of our food-safety feature — its white space. But four UX ideas are worth taking:
Concierge / import onboarding beats the empty form. Labl.it removes the biggest adoption barrier — data entry — by ingesting a menu the customer already has ("Menu Analyser"). We don't need concierge staff, but we already have parse-recipe / parse-receipt infrastructure: an "import/parse my existing menu or supplier list to seed checks/products" onboarding beats an empty checklist. For food safety specifically, this maps to a one-tap "start from a typical SFBB outline" seed (structure only — §4.1, and #440 decision-5 posture).
**Split the author journey from the reuse journey, hard. Their everyday task is 3 taps (category → item → print); authoring is a separate admin surface. For us: recording a due check should be near-zero-friction and must never re-traverse the setup UI.** The Today strip is already this instinct — protect it; don't let "recording today's fridge temp" ever require opening the check editor.
Passive audit logging — capture evidence as a byproduct of normal work. Every print is auto-stamped with product/time/staff initials, giving "audit-ready" evidence with zero extra effort. Our append-only entries already do this in spirit; the lesson is to surface that trail as the selling point ("your diary writes itself, and it's EHO-ready"), and to keep capturing who + when automatically (recorded_by) rather than asking.
Permission split: edit vs. record. Central admin owns the templates; floor staff can only print. Prevents the exact failure mode of a well-meaning staffer editing a check wrong. Worth mirroring once we have multi-user tenants — an owner defines the checks; staff can only record against them (which also protects the append-only integrity).
And one automatic-emphasis lesson that reinforces our posture: Labl.it makes it structurally impossible to print an un-emphasised allergen (bold is automatic, never user-controlled). Same spirit as our "labels are derived, never hand-edited" — a good reminder for any check that touches allergen accuracy.
4. Layout & UX — concrete suggestions, grounded in our design language
This is the section Dave asked for: specific layout/interaction ideas, reconciled with DESIGN-LANGUAGE.md so they fit our app rather than importing GoAudits' paradigm wholesale.
4.0 The one big design decision: "scattered checks" vs a "run"
GoAudits and the FSA diary agree on a paradigm we don't currently use. GoAudits runs an inspection as a paged session (one question per screen, prev/next, sign off at the end). The SFBB diary groups the day into "Opening checks" and "Closing checks" — each a single grouped attestation, signed once. Our model today is individual checks scattered in a Today strip, each recorded on its own.
The suggestion: offer an "opening checks" / "closing checks" run — a single paged (or single-scroll on desktop) flow that walks the due checks in order and ends with one dated, named sign-off ("Our safe methods were followed and effectively supervised today"). This is simultaneously closer to how GoAudits feels and closer to what the FSA diary literally is — and it doesn't replace the Today strip, it's the thing the Today strip's "X due" chip opens into. Individual ad-hoc recording still works for a one-off fridge reading. This is the highest-value layout idea in the whole research, because both reference points independently point at it.
Constraint it must respect (DESIGN-LANGUAGE "nothing may ever tick itself"): a run pre-fills nothing and has no "mark all done" — each item is still an explicit, individually-taken action; the run just puts them in one ordered place with a sign-off at the end. A grouped attestation is fine (SFBB does exactly that); a defaulted tick is not.
4.1 App-level IA — keep it flat; borrow the "urgency-first home"
GoAudits' whole nav is a 3-item bottom tab bar (Home / Audits / Actions) and an urgency-first home (overdue count front-and-centre, tappable straight to the filtered list). We already lead with Today as the home on every screen (#173) and already show a Food safety block there.
Keep our four-view segmented control (Today's checks / History / Corrective actions / HACCP plans) — it's our #369 "select switches the section" pattern and it works.
Borrow the urgency-first treatment: the Today strip's Food safety block should lead with open corrective actions and overdue checks, each a .brow.stack row whose sub is a sentence, with the count as a pill warn (we already do a version of this — tighten it so overdue/ open actions sort to the top).
4.2 The per-check row/card, and the paged runner
From GoAudits' per-question layout (question text → single-select answer row → visible affordance row), adapted to our tokens:
A check in the run = a card (.card) with the check name as the row title, the frequency as the .sub, then the response control:
tick checks → a single explicit "Record ✓" action (not a pre-tick), plus an inline "problem?" affordance (see 4.3).
temperature checks → the numeric input beside the target, using our existing "machine readings shown beside what they'd replace" (#249) treatment: show the target range next to the entry so an out-of-range value is visually obvious the instant it's typed.
Affordances stay visible, not in an overflow (GoAudits lesson): under each check, a small row — Note · Photo (photo later) · Problem/Corrective action. Labelled, not hidden.
Progress: on the paged runner, a simple "3 of 8" counter and prev/next; on desktop a single scroll is fine (GoAudits itself uses upload-only/scroll on its WebApp). Default paged on phones.
Recording a tick check as "problem" (or entering an out-of-range temperature) reveals an inline corrective-action drawer at that check, pre-linked to it, with GoAudits' four fields: what you did / who / priority / due date. This is #182's "detect, don't decide" — the app computes out-of-range; the person decides and describes.
Gate it: on a fail, the correction note is required before the entry commits (the FSA diary's "what did you do?" made non-optional). Respect #182's "don't ask a question the user cannot answer yet" — the gate only fires on an actual fail, never speculatively.
4.4 Corrective-actions list — a quick-filter tab strip
GoAudits' Actions screen is a tab strip (All / Open / Overdue / Closed / In progress) + a filter icon, with overdue elevated everywhere (its own tab, its own count on home). Ours is a flat list today.
Borrow a light filter row on the Corrective actions view (our #451 filters pattern) with at least Open / Overdue / Closed, and sort overdue to the top. A corrective action is a process record (open→closed), plain-editable per DESIGN-LANGUAGE — this is a display change, not a data one.
4.5 The EHO pack (report) layout
GoAudits' PDF is a good skeleton for what #441 already does: cover (business + logo + date range) → summary → per-section results → annotated photos inline → corrective-actions list → signature + timestamp. We print portrait printdoc (the printAllergenMatrix pattern) already.
Add over time: photos inline against their entries (§2.3), the daily sign-off attestations as a signature-style block (§7.2), and the HACCP plan section we already ship. Do not add a score/grade band (§5) — that's the one part of GoAudits' report we deliberately omit.
4.6 Onboarding & the author-vs-reuse split (Labl.it)
Beat the blank checklist: a one-tap "Start from a typical SFBB outline" that inserts neutral, obviously-editable check names only (opening: fridges, staff fit-for-work, areas clean, pest-free, handwashing stocked, hot water, probe working, allergen info accurate; closing: food covered/labelled, use-by discarded, cleaning done, waste out, areas disinfected, washing-up, floors, records complete) — mirrored verbatim from the FSA pack (§6.1) and #440 decision-5 posture (structure, never numbers/hazards).
Protect the reuse path: recording a due check must never open the check editor. The Today strip → run is the reuse journey; the check editor is the author journey. Keep them separate.
4.7 Small ergonomics worth copying
Voice-to-text on the correction note (kitchen hands are busy/dirty); the temperature field already correct; consistent status vocabulary + pill colour reused across Today, History, Actions and the pack (we already have the pill system — just apply it uniformly: warn for due/overdue/open, ok for recorded/closed).
4.8 Calendar / diary view — well-justified, but anchor it on the FSA weekly grid
Researched as its own question (RESEARCH.md §15a, calendar paragraph). The verdict is yes, worth doing — but be precise about which layout, because the evidence points at a specific shape rather than a generic month grid.
The evidence, ranked:
Strongest — official and format-specific. The SFBB diary is a calendar: a week-to-view grid, Monday–Sunday, in 4-week blocks each followed by the 4-weekly review, each day cell holding opening/closing checks + "any problems — what did you do?" + Name/Signed. This is the exact paper artifact our target users already keep and EHOs expect to see — direct precedent for a weekly diary layout, and it dovetails with the opening/closing "run" (§4.0).
Category is mixed on a true calendar grid. SafetyCulture (the leader) does offer a real List / Calendar toggle for scheduled inspections; Lumiform claims one. But GoAudits, Trail, Jolt and Operandio deliberately use daily task lists + due/overdue dashboards, not calendar grids — Trail even markets itself as a digital SFBB-diary replacement yet renders it as a daily list, not a grid. So a month grid is a differentiator, not table stakes.
Highest everyday value is a "what's due" planner, not the month grid. The market treats the agenda/"due today / this week" surface as the daily driver; the month grid is a secondary "see-the-whole-month / prove-no-gaps" affordance.
The suggestion (layered, cheapest-first):
Week view = the SFBB diary page — the anchor. Seven day columns; each day shows its due/ recorded/missed checks and the sign-off state. This is the most defensible default and the natural home for the opening/closing run (§4.0). Reuses our pill colours: warn due/missed, ok recorded.
Month view = a compliance-at-a-glance heatmap — one cell per day, green/amber/red for all-done / partial / missed. This is the visual of the "confidence in management" story an EHO scores (§6.3): prove the diary has no gaps. A read projection over frequency + dated entries — no new data.
Keep the Today strip as the everyday driver (it already is the "what's due" planner the market values most); the calendar is the review/proof surface it opens into.
Hard constraint (DESIGN-LANGUAGE "nothing may ever tick itself", append-only integrity), as refined by §6.6: a calendar is a read/reporting surface. It must never let a user tap an empty past day to silently back-date a tick — that would falsify a due-diligence record. It shows the gap. It may offer to record an honest late entry for that day (§6.6) — but only one that is stamped with the true "recorded now" time, labelled "recorded late", and shown as such on the calendar (a distinct state, not indistinguishable green). A missed Tuesday either stays visibly missed, or becomes a visibly late-recorded Tuesday — never a silently-perfect one.
5. Scoring — a deliberate non-recommendation
GoAudits scores audits (per-option scores, N/A neutralised, overall % + grade bands, RAG dashboards) — but none of it is exposed in their templates; it's proprietary app logic. The FSA's own SFBB has no scoring at all — it's tick + exception + sign-off, and the inspector scores you (FHRS, §6.3). Labl.it has no scoring either.
Recommendation: do not add audit-style % scoring to daily checks. It misrepresents the SFBB model (pass-by-doing) and invites a "we scored 92% so we're compliant" misread — the exact compliance-claim overreach our posture avoids. If scoring ever appears, it belongs only to a clearly-labelled FHRS self-assessment ("how an EHO might see your records," an estimate, never a verdict) — a separate decision, not part of the daily-checks feature.
6. What the FSA actually requires — and how it validates our design
(Primary-source detail and verbatim quotes in RESEARCH.md §14b — the full 100-page SFBB caterers pack + CookSafe were extracted directly this pass, closing §14/§14a's gaps.)
6.1 The daily diary is exception-based, with a signed sign-off — now confirmed verbatim
Each day = a tick pair (Opening / Closing) + a free-text "Any problems or changes – what did you do?" + Name / Signed under "Our safe methods were followed and effectively supervised today." Validates our simple-tick + separate-corrections design, and the verbatim opening/closing lists (incl. the FSA's own "Allergen information is accurate for all items on sale" opening check) are the source for the §4.6 seed outline. That allergen check ties a legal daily food-safety check straight into our PPDS/label engine — a genuine product hook.
6.2 The numbers now have a primary source (England/Wales/NI)
Verbatim from the SFBB pack + GOV.UK: chilled 8°C (legal) / fridge 5°C (recommended); freezer −18°C; cook 70°C/2min + the equivalence table (80/6s, 75/30s, 70/2min, 65/10min, 60/45min); hot-hold 63°C; cold-display 4-hour rule; probe calibration iced −1 to 1°C, boiling 99 to 101°C. England publishes no numeric cooling/reheat figure. Scotland/NI have a statutory 82°C reheat; England/Wales don't — which is exactly why "store the figure as data, never hard-code" is right.
6.3 FHRS — what our users are judged on, and where our feature moves the needle
An EHO scores three areas (lower = better) and the rating is capped by the worst area, not the average: Food hygiene & safety / Structure / Confidence in management. Confidence in management is the area our feature most directly improves — complete, honest, up-to-date, non-back-fillable records are precisely the evidence of "confidence in standards being maintained in future." This is the commercial argument for the feature, and the copy framing (§7.7).
6.4 Record-keeping / retention — the honest answer
A defensible record needs what / who / when / corrective action / periodic review — all of which our schema captures. No statutory retention period exists (SFBB states none; GB law sets none — duty is "kept up to date and available for inspection"). The app must not assert a fixed legal retention figure.
6.5 Non-operating days — records track operating, not the calendar (VERIFIED, RESEARCH.md §14c)
The FSA does not require a record on a day the business didn't trade. The SFBB diary is built around opening and closing the business ("sign the diary every day… opening and closing checks have been done"), and the legal basis (Reg. (EC) 852/2004 Art. 5, retained) is proportionate to the nature and size of the business — so a seasonal/occasional/part-time business (market stall, event caterer, farm shop closed some days) keeps records for the days it operates, full stop. EHO concern is gaps on days you operated, not blank non-trading days.
Design consequence: a non-operating day is "not required", not "missed". The app must know when a business operates — via a trading-days pattern (+ ad-hoc closures) or an explicit "open today" action (which also starts the opening-checks run, §4.0) — so the calendar shows closed days as neutral, never red, and "missed" only ever means an operating day with checks not recorded. Forcing a daily cadence on a 3-days-a-week business is our bug, and actively misleading as evidence (it paints a compliant business negligent).
The one caveat:time-based controls on stored perishable stock (fridge/freezer temps, scheduled cleaning) attach to stock existing, not to trading — so they can still be live on a closed day. Distinguish "closed & empty" (nothing due) from "closed but holding stock" (stock-temperature checks may still apply).
6.6 Late entries — contemporaneous is the ideal; honest late is allowed; back-dating is falsification (VERIFIED, RESEARCH.md §14c)
The sharper principle behind "no back-filling": the integrity requirement is "no falsified records", not "no late records". UK good-documentation practice (MHRA ALCOA+ — the authoritative UK articulation; food EHO/consultant sources say the same) is explicit:
Records should be contemporaneous (made at the time) — the ideal, and what SFBB assumes.
A late entry is permitted when it (a) records the true date/time it was actually written, (b) also states the date it attests to, (c) is visibly flagged as a late entry (optionally with a reason), and (d) never overwrites or disguises itself as an original contemporaneous record.
Back-dating is falsification — fatal to the Food Safety Act 1990 s.21 due-diligence defence, and EHOs reportedly spot back-filled paper diaries "in the first 15 minutes." An honestly-marked late entry is far stronger evidence than a back-dated one (it may read as a minor process weakness; a back-dated one reads as fraud).
Design consequence: never allow silent back-dating. When a user records after the fact, capture and display both timestamps — the tamper-evident system entry time (immutable, our existing created_at append-only truth) and a new "attests-to" date — and render the entry as "recorded late" everywhere it appears (History, calendar, EHO pack). This removes the barrier for the did-the-check-forgot-to-log business and is exactly what protects their due-diligence defence.
Design principles this research surfaced (proposed — graduate to DESIGN-LANGUAGE.md when built)
Two refinements/additions to our compliance-record design language, both now primary-source-backed:
Records track operating, not the calendar (§6.5). "Due" and "missed" are defined relative to the business's operating days, never the raw calendar. A non-operating day is a first-class closed state, not an absence. (New principle.)
Never falsifiable, not never-late (§6.6). Refines the append-only rule and the §4.8 calendar constraint: the app forbids back-dating, not late recording. A late entry is allowed iff it is stamped with the true entry time, carries an attests-to date, and is visibly labelled late. (Refines "Append-only records with a correction trail" + "nothing may ever tick itself" — a late entry is still one explicit human action, never a default or a bulk fill.)
7. Recommendations (for Dave to triage into issues — none built)
Organised into waves; nothing here has been built. The GitHub epic mirrors this structure.
Wave 0 — foundational (unblock the rest; encode the two design principles):
F1. Operating-days / trading-pattern awareness (§6.5) — checks due only on operating days; closed days a neutral state, never "missed"; distinguish closed-but-holding-stock. Unblocks a truthful calendar.
F2. Honest late attestation (§6.6) — record after the fact with immutable entry time + an attests-to date + a visible "recorded late" marker; never back-date. Data-model + integrity.
Wave 1 — quick wins on existing machinery:
Q1. Inline corrective action, note gated on a fail (§2.1, §2.2, §4.3) + Open/Overdue/Closed filter tabs on the Corrective actions view (§4.4). Highest value-to-effort.
Q2. Daily "safe methods followed & supervised today" signed sign-off (§4.5, §6.1) — append-only attestation, verbatim SFBB sentence; strengthens the EHO pack.
Wave 2 — the layout payoff:
L1. Opening-checks / closing-checks "run" (§4.0) — day grouped into one ordered flow ending in a sign-off, opened from the Today chip. Matches both GoAudits and the FSA diary; no auto-ticking.
L2. Week (SFBB-diary) + month heatmap calendar view (§4.8) — read-only projection over frequency + dated entries; states closed / on-day / late / partial / missed. Depends on F1+F2.
Wave 3 — depth & evidence:
D1. "Start from a typical SFBB outline" seed (§4.6) — structure-only check names, verbatim from the FSA pack; #440 decision-5 posture.
D2. Photo evidence on check entries + corrective-action resolutions (§2.3, §4.5) — real storage cost; high compliance value in the pack.
D3. Guided 4-weekly review checklist (RESEARCH.md §14b) using the verbatim FSA review items.
P1. Positioning & docs — FHRS-framed value copy (§6.3), "your diary writes itself, EHO-ready" passive-logging framing (§3.3), and sourcing our reference numbers to the FSA pack (§6.2).
P2. Later / multi-user — permission split (owner authors checks, staff only record — Labl.it, §3.4); voice-to-text on notes (§4.7).
Explicitly not recommended: audit-style % scoring on daily checks (§5); seeded hazard/limit content (unchanged #440 posture); hardware/device dependency (Labl.it's model — our browser-based, print-anywhere approach is a differentiator to keep).
7. Accounts / simple bookkeeping for home bakers (researched 18 Jul 2026)
(A second §7 by historical accident — citations of "RESEARCH §7" elsewhere may mean either this or Open questions above. Kept unrenumbered so no existing citation breaks.)
Context: Dave wants an Accounts function — record cost of purchases and revenue from customer orders, on a monthly basis. Deliberately not complex given the audience.
What the market/best practice says:
Keep it simple — no double-entry. Cottage-food / home-baker guidance is unanimous that simple income + expense tracking plus a profit-and-loss report at tax time is enough; double-entry bookkeeping is overkill. Wave (free) or even a spreadsheet is considered sufficient for most. (BakingSubs, FreshBooks.)
Cash basis is the right default (and now HMRC's default). From 2024/25 the cash basis (record income when money is received, expenses when paid) is the default for UK sole traders, eligible up to £150k turnover ⚠️ corrected 1 Aug 2026 (#341): the £150k/£300k turnover limits were abolished from 2024-25 — cash basis is now available to any size of unincorporated business, and accruals is the opt-out (a tick-box on the return). It aligns taxable profit with actual cash flow and is the simplest mental model for a baker. (CWABC, LITRG; correction ICAEW/ ACCA.) → We should build cash basis. (Also superseded in part: the MTD dates bullet below — see §13 for the confirmed £50k/£30k/£20k phase-in.)
Trading allowance £1,000. If gross trading income in the tax year is ≤ £1,000, a sole trader needn't register/report; above that they can deduct the £1,000 flat allowance instead of actual expenses. Useful context to surface, not something we must enforce. (LITRG.)
The three jobs of the software: (1) track income, (2) track expenses, (3) produce a simple P&L. Everything else is a bonus. (BakingSubs, FreshBooks.)
COGS is an optional advanced layer. Craftybase-style tools compute cost-of-goods-sold per sale from the recipe's material costs and show per-order profit; bakeries target ~40–60% margins. For a simple v1, total purchases as expenses (cash basis) is enough; per-sale COGS can come later and we already have the ingredient/recipe cost machinery to support it.
Capital items: under cash basis, most equipment is just expensed in the month paid (only cars are special) — so a flat "expense in the month you pay" model is correct. (CWABC.)
MTD for Income Tax starts Apr 2026 for >£50k income (>£30k Apr 2027) — most home bakers are well under, so not a near-term constraint, but a reason to keep clean monthly totals. (RRAccountants.)
Implications for our build:
Model as a simple transactions ledger: dated entries, each either income (money in) or expense (money out), with an amount, a category, a note, and an optional link (to a customer order for income, or an ingredient/purchase for an expense).
Monthly P&L view: pick a month → total income − total expenses = net profit, with a category breakdown. Cash basis (count it in the month the money moved).
Revenue from orders needs order prices. Orders (#31) currently store no price/total, so to auto-derive revenue we must add a price to orders (this is backlog #61). Decision needed: derive income from orders vs. manual income entries vs. both.
Reuse existing cost data where sensible (sourcing pack costs, recipe/bake costs) but a purchase is a dated cash event, distinct from the reference unit-cost we already hold.
Sources: see list below (accounting section).
Capability
Last updated 19 Jul 2026
8. Receipt OCR / parsing (researched 19 Jul 2026)
The receipt scanner (#67) uses on-device Tesseract (tesseract.js v5, no upload). Real-world reads were sometimes poor — the worst failure being two physically separate receipt rows collapsing into one item line (e.g. an Aldi receipt merged 496633 BAKING NUTS 1.09 A and 416555 BUTTER UNSALT 250G 7.96 A into a single line, taking only the last price). What the literature says works:
Reconstruct rows from word bounding boxes, not from the flat text. Tesseract returns every word with a bbox (x0/y0/x1/y1) and a confidence. Grouping words by their vertical centre into rows, then sorting each row left-to-right, recovers the true physical lines even when Tesseract's own text output merged them. This is repeatedly cited as the single most effective fix for receipts (they are relational/tabular documents, not prose). Implemented as linesFromOcrData() (19 Jul), with a safe fallback to data.text when word boxes aren't present.
Page-seg mode matters:--psm 4 ("a single column of text of variable sizes") suits receipts; we deliberately did not switch to a Tesseract worker to set it — an earlier worker rewrite (v0.24.1) broke the engine to 0 lines (see HANDOVER §4), and the proven path is Tesseract.recognize.
Drop very low-confidence words to cut noise before row assembly.
Reference/transaction numbers vary widely (added 19 Jul, v0.36.2). The old parser only read Tesco's 4-group 3RJC-1TRE-M058-YXXL code, so most receipts came through with a blank reference. Receipts label it many ways — Receipt No, Transaction, Trans, Txn, Invoice, Order, Reference/Ref (a genuine receipt number), and Auth Code, Sequence/Seq, EFT (a payment identifier) — or print an unlabelled code (Tesco's 4-group; Aldi's slash-grouped C456/014/807). parseReceiptReference now tries labelled values first (receipt-family before payment-family), then the 4-group, then the slash code, rejecting dates, times, prices and masked card numbers. Don't take Terminal/Merchant ID (not transaction-unique) or the card number.
Parse defensively: even with good row separation, a line can still carry two price + tax-letter tokens; the parser now splits a row into one item per price token (parseReceipt/splitReceiptRow), handles Aldi's N x unit multiplier rows (qty/unit attach to the following item), and strips the leading 4–14 digit item code from the printed name. This means the merge is fixed at parse time too, independent of OCR quality.
Summary/total lines survive OCR damage and leak in as items (added 19 Jul, v0.36.3). Seen live: a "Total: GBP 14.14" read as otal: GBP and a "…Items/Sales" total read as Itens Sales both got added as £14.14 items. Exact-word summary matching can't catch a mangled keyword. Two-part defence: (a) OCR-tolerant summary terms (otal for a T-dropped Total, gbp, sales?); (b) a post-parse sweep that drops any parsed item whose description still reads as a summary. Kept it safe — baking-ingredient names don't contain Total/GBP/VAT/Card/etc, so real items aren't at risk — and a price-less pending line no longer flushes on a summary line (which had let an address line leak once GBP counted as a summary word). A fuzzier amount-based total detector (a line whose amount ≈ the sum of the others) was considered but shelved as riskier — it can drop a real single-item line and is unreliable when several total lines repeat the same figure.
Heavier options exist but don't fit our constraints: cloud OCR (Google Vision, Textract) and LLM-based extractors read receipts far better, but the product's promise is on-device, no upload, photo never stored — so we stay with local Tesseract + smarter parsing. > Corrected 27 Jul 2026 — we did the opposite, and the old promise is dead. The literature was > right that LLM extractors read receipts far better, and that is now the primary path: > parse-receipt (#168, Mobile M2) sends image_base64 to an edge function that reads the photo > with Claude vision. On-device Tesseract survives only as the fallback. Receipt images are also > stored now, in a private bucket with 6-year default retention (#120). > > So "on-device, no upload, photo never stored" is false on all three counts and must never be > repeated in marketing copy — it nearly was, in the 27 Jul commercial report, which is corrected. > The honest privacy claim is UK data residency: the Supabase project is eu-west-2 (London). > The Anthropic call is a US transfer requiring a documented transfer mechanism, and Anthropic is a > sub-processor who must appear on the published list.
Sources: see receipt-OCR entries in the list below.
Capability
Last updated 28 Jul 2026
8b. Dext as a reference for expense capture (#146, researched 28 Jul 2026)
Timeboxed spike, sequenced deliberately before #182 and #183 because what the category leader actually ships should shape both. Dext is the reference point for receipt capture in UK bookkeeping: ~700,000 businesses, Xero Small Business App of the Year 2024.
⚠️ Evidence caveat, stated up front.dext.com and help.dext.com both return 403 through the agent proxy, so none of their own pages could be read directly. Everything below comes from search-result extracts and third-party reviews. It is good enough to shape a design and not good enough to copy a spec from — anything marked below as unconfirmed needs a look at the real app before it drives a build.
What they do that bears on #182 (focused capture)
No evidence of live edge detection or auto-capture. Repeated searching surfaced capture modes and a crop tool, never an auto-detecting viewfinder. Unconfirmed — absence in search results is not absence in the product — but it is a genuine signal that the market leader has not found live outline detection necessary, which materially lowers the priority of the hard half of #182.
Crop is POST-capture, via a crop icon on the review screen. That is precisely the fallback #182 offers ("if this can't be done then at least an ability to crop and adjust before saving"), and it is what the leader ships as its primary answer.
Three capture modes, and this is the most useful finding of the spike:
Single — one receipt, one page
Multiple — several separate receipts in one go (up to 50)
Combine — one document spanning several photos, explicitly for long receipts and multi-page invoices (up to 50)
Combine is how they solve the long-receipt problem — by structure, not by optics. A till roll that will not fit one frame gets photographed in overlapping pieces and treated as one document. That is far cheaper than a clever camera and solves the case a supermarket shop actually produces.
Implication for #182: re-scope it. The expensive part (live outline detection at capture time) has no evidence behind it; the cheap parts — post-capture crop, and a Combine-style multi-photo single document — carry most of the value. Worth re-writing the issue before building. Done 28 Jul 2026: #182 rewritten to exactly that scope, with live outline detection descoped (not rejected) and the original request preserved at the foot of the issue. (Shipped: crop in v0.81.0, Combine in v0.82.0.)
What they do that bears on #183 (email-to-scan)
Each account gets a unique inbound email address to forward receipts and invoices to. This is the standard pattern and is exactly what #183 describes, so #183's shape is validated.
"Fetch" pulls invoices directly from supplier portals — a category beyond #183 entirely, and a reminder that email-in is the middle rung: photo → email → direct supplier connection.
Implication for #183: the unique-address model is the right one, and the security questions it raises (who may send to that address, how spoofing is handled, what happens to attachments that are not receipts) are the real work — not the plumbing. Unchanged conclusion: this needs Dave's infrastructure and security decisions before it is buildable. Done 28 Jul 2026: #183 updated with six named decisions (sender authorisation, spoofing / SPF-DKIM-DMARC, address secrecy and rotation, non-receipt attachments, failure visibility, inbound provider and UK data residency) and labelled security — it opens a write path into tenant data authenticated by something other than a signed-in session. (Shipped in v0.80.0; the decisions are §8d and the Worker README.)
What they do that we could extend cheaply
Supplier Rules — per recurring supplier, default the category, tax code and payment method. "Always categorise Costa Coffee as Meals & Entertainment." We already have supplier profiles (#101) that resolve a merchant name from a receipt ("TESCO STORES 3294" → your "Tesco"). Extending those to carry a default expense category is a small change on top of something that exists, and it attacks the same zero-typing goal M1 is for. Raised as #245 on 28 Jul 2026 (M1). Reading the code sharpened the case: #69 catalog matching already sets a category per line, so the real gap is the unmatched line, which falls back to "Supplier products" unconditionally — wrong on every line for a packaging or equipment supplier, and wrong most often during onboarding, when the catalog is empty so nothing matches. Tax code and payment method were deliberately left out: we have no such fields, and inventing them to complete someone else's feature would be building for an accounting workflow nobody has asked us for. (Shipped in v0.76.0.)
Line-item extraction — individual items, not just totals. We already do this, which is worth recording: on the capability that matters most for costing, we are not behind.
The accuracy claim, read honestly
Dext markets "up to 99.9% accuracy" for supplier, date, totals, tax and line items, "verified across millions of documents". Treat as marketing, not a benchmark: self-reported, no published methodology, no error taxonomy, and "up to" is doing real work. It is not a bar we have been shown to fall short of, and it should not be quoted back at us as one. Our own evidence base for parsing quality remains §8a, built from real receipts.
8d. Inbound email providers and UK data residency (#183, researched 28 Jul 2026)
Decision 6 of #183 — the only one Dave left open. Brief: keep UK data residency if we can.
The framing that resolves it: residency is about STORAGE, not about every processor
The claim in §8 is "UK data residency — Supabase is eu-west-2 (London)", and it replaced a false claim that receipts never leave the device. It has never meant "no byte ever leaves the UK" — Anthropic already processes every scanned receipt (#168) and is named as a sub-processor for exactly that reason.
So the question is not "can mail be received in the UK" but "where does a durable copy come to rest?" If the inbound provider is configured to retain nothing — parse, hand the attachment to Supabase in London, keep no copy — the only lasting copy is in the UK and the claim stands, with a line added to the sub-processor list.
Legally, EU hosting is fine either way. UK adequacy regulations cover the EEA, and the European Commission renewed its UK adequacy decisions on 19 Dec 2025, so data flows both directions with no additional safeguards. This is a claim-accuracy question, not a lawfulness one — but the ICO still expects proportionate checks on the recipient, which is an argument for a provider with a real EU region rather than a global one that merely might keep data close.
The candidates
Provider
Inbound region
Verdict
Mailgun
Real EU region (api.eu.mailgun.net); strong inbound routing/parsing
First choice
AWS SES
Inbound in 3 regions only: us-east-1, us-west-2, eu-west-1 (Ireland)
Viable, Ireland
Cloudflare Email Workers
Global edge; no contractual guarantee below Enterprise, but stores nothing
✅ CHOSEN 29 Jul — see §8d-1
Postmark
US only, no EU plans
❌ Ruled out
SendGrid
US inbound parse
❌ Ruled out
⚠️ AWS SES CANNOT RECEIVE MAIL IN LONDON. From AWS's own documentation source: email receiving is supported in us-east-1, us-west-2 and eu-west-1 only, and Europe (London) appears explicitly on the "not supported" list alongside Frankfurt, Paris, Milan and Stockholm. This is the single most decision-relevant fact here — the obvious "just use SES in the same region as Supabase" answer does not exist. (One 2026 third-party guide referenced an inbound-smtp.eu-west-2 endpoint, which contradicts the docs; the AWS source was read via its GitHub docs mirror, which can lag. Confirm in the SES console before relying on it either way.)
Cloudflare is the interesting one. It is already in the stack — support@provenbatch.co.uk runs on Email Routing (free, #213) — and Email Workers process inbound mail through an email() handler that streams rather than stores, which is exactly the "retain nothing" shape the framing above wants. If processing is transient and the attachment goes straight to Supabase, residency of stored data is preserved by construction and no second vendor enters the stack. But Cloudflare is a global edge network and their residency guarantees are a paid add-on (Data Localization Suite), and their docs 403 through the proxy so this could not be confirmed. Worth 30 minutes before defaulting to Mailgun, because if it holds it is clearly the best answer.
⚠️ This was answered on 29 Jul 2026 — see §8d-1 below. Short version: the guarantee does NOT exist at our tier, and we took Cloudflare anyway, for reasons that are written down rather than assumed.
Recommendation
Check Cloudflare Email Workers first — already in the stack, free, streams rather than stores. If it retains nothing, take it.
Otherwise Mailgun EU — a real EU region and the strongest inbound parsing of the three.
AWS SES eu-west-1 as the cheap fallback, noting it puts an Irish S3 bucket in the path unless the object is moved to Supabase immediately and deleted.
Rule out Postmark and SendGrid on residency alone.
Capability
Last updated 29 Jul 2026
8d-1. The Cloudflare answer, and what we actually shipped (#183, 29 Jul 2026)
The 30 minutes above were spent. The answer is not the comfortable one, and the distinction matters more than the verdict.
What is true
Data Localization Suite is Enterprise-only. There is no contractual residency guarantee available on our plan, at any price short of Enterprise. The "needs proof" verdict above was looking for a guarantee that does not exist for us.
But Email Workers genuinely stream. Mail arrives at an email() handler; nothing is persisted unless the Worker opts in to Cloudflare-side storage — KV, Queues, D1, R2, Durable Objects. Storing an inbound message in KV is a documented pattern, but it is a pattern you choose, not the default.
So both halves of the earlier hope were right and wrong at once: the behaviour is what we wanted, the guarantee is not available.
What we decided, and the honest framing
We took Cloudflare and shipped it in v0.79.1. The reasoning, stated plainly so nobody later mistakes it for something stronger:
The residency position rests on "no durable copy comes to rest at Cloudflare" — which is the framing §8d established and which is satisfied by construction, because the Worker declares no storage bindings at all. It is a framing, not a guarantee. If a customer ever asks for a contractual residency commitment, the honest answer is that we do not have one from Cloudflare, and the remedy is Mailgun EU.
That remedy stays cheap on purpose: the Worker is a thin adapter (~200 lines), and everything around it — the token, the allowlist, the source='email' row, the review queue — is provider-independent. Swapping providers is a rewrite of one file, not an architecture change.
🔴 The invariant, and why it is enforced three times
The whole argument collapses the moment someone adds a storage binding "just for a cache" or "just for rate limiting". That is a one-line change that silently retires a privacy claim, so it is guarded at three levels, deliberately redundant:
A test can be deleted by the same person adding the binding
The deploy token
Scoped to Workers Scripts + Account Settings only — it cannot create a KV namespace or R2 bucket
—
The third is the real one: it makes the invariant unforgeable rather than merely asserted. If a future change genuinely needs Cloudflare-side storage, widening that token is the moment to stop and re-read this section, rather than discovering the claim was quietly false months later.
Rate limiting is counted in Postgres for exactly this reason — the obvious place for it would have been KV, and that would have cost the argument.
The sub-processor consequence is better than expected
§8d assumed taking Cloudflare would add "a line to the sub-processor list". It adds nothing. D-12 already lists Cloudflare (hosting + DNS, corrected 29 Jul when Netlify was retired by D-20/#251), and D-20 moves the app, staging and both marketing sites onto Cloudflare Workers anyway.
So the choice that looked like it needed a residency proof turns out to add no new vendor, no new disclosure and no new cost — which is a materially stronger position than the Mailgun alternative, where a genuine second processor would have entered the stack. Worth weighing against the missing guarantee rather than treating the two as separate questions.
One more thing worth recording: DMARC's limit
Not residency, but it surfaced building this and belongs with the evidence base. DMARC authenticates the domain in From:, not the local part. It proves a message came from gmail.com; it does not prove it came from that Gmail user. For the big providers the distinction is academic — they do not let one account send as another. On a small self-run domain where any user can send as any address, allowlisting one person there effectively trusts everyone on it. That is a property of email, not a fixable defect, and it caps how much the #183 allowlist can promise.
Whichever is chosen: configure retention to zero, and add the provider to the published sub-processor list beside Anthropic.
⚠️ Evidence caveat. Only the AWS regions list was read from a primary source (AWS's own docs mirror on GitHub). developers.cloudflare.com, docs.aws.amazon.com and the Mailgun docs all 403 through the agent proxy, so the Cloudflare and Mailgun findings are search-extract evidence. Good enough to shortlist; not good enough to sign a contract on.
8c. BarcodeDetector on Safari — the WASM fallback is NOT dead weight (#170, checked 28 Jul 2026)
#170 says "BarcodeDetector where available + lazy-loaded WASM fallback (verify Safari support at build time)", and HANDOVER carried the hope that if Safari had shipped it the fallback could be dropped before it was ever written. It has not shipped. The fallback is the primary path on iOS, not a contingency, and that should be assumed from the start of the work rather than discovered halfway through.
What the evidence says, consistently across several independent sources:
Not implemented in WebKit, and since every browser on iOS is WebKit, that means no iOS browser has it — not Safari, not Chrome for iOS, not Firefox for iOS.
There is a partial implementation behind a feature flag, off by default. Guidance is explicitly to treat it as unavailable, because no ordinary user has toggled it.
WebKit bug 281848 — "Shape Detection API doesn't work on iOS" — reports it broken since the iOS 18 update, i.e. the flagged implementation is not merely absent but faulty.
Chromium-only in practice (as of early 2025 reporting). Firefox has a request ticket that has not been worked on.
Not a W3C standard and not on the Standards Track — the Shape Detection APIs remain a draft.
Two consequences for how #170 gets built:
The WASM decoder is the iOS path, so it must be the well-tested one. Dave's own devices are iPhones and both known PWA installs are phones — so on the hardware this feature exists for, BarcodeDetector will essentially never be the code that runs. Building it "native first, fallback later" would test the branch that matters least.
⚠️ Feature-detect by CAPABILITY, not by existence.'BarcodeDetector' in window is not safe here: a user who has enabled the experimental flag would pass that check and then hit the broken iOS implementation. The detection must actually attempt a decode against a known image and fall through to WASM on failure or exception. An existence check would produce a scanner that silently never reads anything — precisely this project's recurring failure shape.
Re-check trigger: worth asking again only if WebKit's standards position changes or a WWDC release notes it. Nothing here is time-sensitive — the answer has been stable for years.
⚠️ Evidence caveat.developer.mozilla.org and caniuse.com both 403 through the agent proxy, as does bugs.webkit.org, so none of the primary compatibility tables could be read directly. This is search-extract evidence, consistent across sources but not first-hand. A five-minute check on a real iPhone would settle it beyond doubt, and that check is anyway needed before #170 ships.
8a. Supplier receipt formats — living record (started 19 Jul 2026)
Why this exists: rather than hand-build brittle per-supplier parsers now (premature — the generic parseReceipt already handles the observed variation, and we have no evidence of which suppliers it fails on), we keep a record of each supplier's receipt layout as real receipts come in. This is the evidence base: if a specific supplier keeps parsing badly, this tells us exactly what a targeted tweak would need to do. (13 Aug 2026: a plan to make this loop systematic — DB-backed chain profiles feeding the AI parser, cross-tenant quality telemetry, admin-console curation — is drafted in outputs/docs/receipt-parsing-knowledge-base-plan.md, awaiting Dave's decisions.) Add a supplier's row the first time you scan one of their receipts; note anything the generic parser got wrong. Fields: item-line shape · where the qty×unit sits · tax letters · date format · a reference/fingerprint unique to that chain · footer keywords · parser notes.
Aldi (id 3) — observed 19 Jul 2026 (real receipt, Aylesbury store).
Item line: <6-digit code> <NAME> <price> <taxletter> e.g. 417009 EGGS LARGE FR 12PK 8.67 A.
Multi-buys: a separate line ABOVE the item, N x <unit> e.g. 3 x 2.89 (3 × £2.89 = £8.67). Single items have no such line.
Tax letter: single A (zero-rated shown as A 00.0% Net). Date: dd.mm.yy in the payment block (*9509 C456/014/807 14.07.26 08:55).
Reference / transaction number = the Cnnn/nnn/nnn slash code on the payment line (e.g. C456/014/807). Settled 20 Jul after iterating: took the slash code, then the *NNNN, then the *NNNN+slash combined — all imperfect/fiddly — so per Dave we simply take the slash code. NOT the *NNNN, NOT the Auth Code. parseReceiptReference reads the slash code ahead of payment labels, skipping dd/mm/yyyy dates (v0.36.6).
Header ALDI STORES + address. Footer: Total, N Items, Card Sales, Goods:, A 00.0% Net … Vat.
Repeated-unit lines: Aldi prints each unit of a product on its own line with the same code + name + price (not N × name) — e.g. 7 separate 387457 SUGAR SOFT 500G 1.09 A lines. It also uses a N x unit multi-buy line above one item (e.g. 12 x 1.45 → one 12-pack). The parser handles both: N x unit sets qty on the next item; identical repeated lines are aggregated by name+unit price into one line with the summed qty (v0.37.0, backlog #114).
Promotions/discounts print as a Promotion header + BUY 2 … FOR £x -0.19 lines (negative amount) — dropped (negative-amount lines are skipped). VAT-rate lines (A 00.0% Net … Vat) and the *ref line (*0132 C456/002/012 18.06.26 13:11) are dropped as footer/date lines.
Parser status: handles the long/dense receipt — code stripped, N x unit → qty, repeats aggregated, promos/footer dropped, slash-code reference captured. Remaining risk is raw OCR quality on a long, crumpled thermal receipt (garbled names won't aggregate with their clean twins).
Tesco (id 1) — observed from an earlier real photo + test fixture.
Item line: usually <qty> <NAME> <price> (name-led, £-prefixed prices common), e.g. 2 Tesco Baking Spread 500g £3.00; sometimes the price OCRs onto its own line below.
Unit price printed as a separate £x.xx each line (ignored by the parser).
Reference: 4 hyphen groups by the barcode, e.g. 3RJC-1TRE-M058-YXXL (parseReceiptReference). Date dd/mm/yyyy. Footer: Subtotal, TOTAL, Clubcard, Card.
Parser status: handled well; watch for Clubcard/points lines (already in the summary filter).
Asda (id 18) — real evidence via #148 (22 Jul 2026); layout notes web-researched.
Tax letter prints AFTER the amount: V = standard-rated, D = zero-rated (groceries/fresh), M = kiosk/tobacco standard, blank = exempt (postage/lottery). The parser's price regex already tolerates one trailing letter, so these don't corrupt the amount.
Reference = the TC transaction code — a long grouped run of digits near the footer, e.g. TC: 3757 5108 7387 3809 4787 0 (REF_ASDA_TC, grouping preserved). VAT no. GB 362 0127 92 at the bottom. Date dd/mm/yyyy.
Parser status: TC captured (v0.47.0 #148); profile prefers it first (v0.48.0 #101).
Web-researched core notes (v0.48.0 #101) — not yet confirmed against a real receipt
These seed each supplier's profile so parsing is supplier-aware the moment the user picks the shop (#150). Treat as provisional until a real receipt lands — then correct in place and promote to a dated "observed" block above. Encoded in the app as SUPPLIER_RECEIPT_PROFILES (keyed by chain name, never id); this record is the human-readable mirror.
Sainsbury's (id 5): name-led item lines; Nectar points footer; transaction number is labelled (generic Tier-1 label reads it) — no distinctive unlabelled fingerprint confirmed, so the profile forces none. Date dd/mm/yyyy.
Costco (id 17): item-number-led lines <item#> <SHORT CAPS DESC> <price> <A|E> (A = taxable, E = exempt — shopper convention, varies by location); member number in the header; the 4–14 digit leading-code rule already strips the item number. Transaction/receipt number is labelled/barcoded.
B&M (id 16) and The Range (id 4): standard thermal; labelled transaction/receipt number (generic Tier-1); VAT number in the footer. No distinctive fingerprint yet — capture on first scan.
Online / trade suppliers — Amazon (10), Vanilla Mart (11), Sephra (6), Ingenious Edibles (8), Rainbow Dust (9): these send invoices / order confirmations (tabular line items, VAT breakdown, a labelled Order/Invoice number the generic Tier-1 reads), not thermal receipts. Flagged invoice: true; item parsing may need an invoice path when a real one is captured.
Still fully uncaptured (no evidence, no research beyond "it's a till receipt"): none of the above block — every current supplier now has at least researched notes. Homemade (7) isn't a real receipt. When any real receipt arrives, replace its provisional block with a dated observed one.
Irish (EUR) receipts — #1521, 3 Oct 2026: no real Irish receipt has been recorded yet. The euro path (parse-receipt's EUR block, the on-device fallback) is built against the shapes the issue names — € before or after the amount (€2.45, 2,45 €, 2,45€), comma decimals, N x €unit multi-buy rows, VAT rate letters keyed at the foot, day-first dates — as constructed fixtures in outputs/tests/ireland1521.js, not as observed formats. When the first real Irish receipt is scanned (Dunnes, SuperValu, Tesco IE, Aldi IE, Lidl IE — the chain fingerprints are their own issue), record its row here as for the UK chains above and correct the fixtures to match. Grouped amounts (1.234,56, 1,234.56, also with a glued VAT letter, 1.234,56A) are read as one number on the euro fallback up to 9.999,99 (#3106 reviews). A space-grouped amount (1 234,56, ASCII or no-break/thin space, which may be a quantity and a price: 12 100,50) and any amount of five or more integer digits (10.500,00, 12345,67, which the price reader would otherwise read as its last four) are kept with a blank amount for the user to fill, never guessed. GBP (#3114, 4 Oct 2026): the UK fallback had the same gaps (£1,234.56 read as 234.56, 12345.67 as 2345.67). It now joins a comma-grouped amount (1,234.56 → 1234.56; a UK receipt never groups with a dot) and keeps a five-or-more-digit or malformed amount (1,234,56, 1.234.56, 12.34,56, 1.234,56) blank. An ordinary space is not GBP grouping (stand-in review of #3151): UK receipts do not group pounds with a space, but wholesale rows print "description qty price" (Wholesale Eggs 30 120.00 = £120.00), which must keep reading correctly, so only typeset no-break/narrow/thin-space grouping is kept blank. The cost, accepted: a genuinely space-grouped 1 234.56 would read as 234.56. The euro fallback keeps a malformed one blank too since #3142 (4 Oct 2026). Also 4 Oct (#3153, #3155): a GBP refund printed -£1.00 is skipped as a refund (the £ used to hide it, so it became a +£1.00 item), a minus glued between digits (Item 12-1,234.56) is an item code, not a refund, and a dd.mm.yy date inside an item line with 6+ letters (Wedding Cake 01.10.26 25.00) is lifted out before the line is read, on both currencies — the date is never the price and never its tail. A dated line whose remaining amount is time-shaped (0.00–23.59, one- or two-digit hour: Served by Sarah 04.10.26 14.32 as UK tills print it, and Birthday Cake 01.10.26 4.50, which cannot be told from 4.50 am) is kept blank, never priced at the time (stand-in reviews of #3161; Dave's rule). #3168 / #3157 (4 Oct 2026): the euro refund test now matches GBP's (any number of integer digits, the U+2212 minus), and the euro unreadable-amount rule's left boundary is (^|\D), so Mehl A,12345,67 is blank, not its tail. An ASCII-space 1 234.56 on GBP stays read as 234.56 by decision: it is indistinguishable from "qty 1 at £234.56", the other reading of the same characters — only a real receipt grouping pounds with a space would change that. A dated item start that is superseded by a priced line is kept as a blank item; an undated one is still dropped, because on the real receipts in receipt.js those are address and header lines.
Sources: real Aldi receipt (Dave, 14 Jul 2026); real Tesco receipt + receipt.js fixtures; real Asda TC (Dave, #148); web research on Tesco/Asda VAT-code letters and Costco receipt layout (22 Jul 2026, see §Sources — smallbusinessowneradvice.co.uk Tesco/Asda VAT FAQs, pricematcher Costco guide).
Capability
Last updated 1 Aug 2026
13. Accounting capability: Self Assessment, MTD, and the FreeAgent benchmark (#341, researched 1 Aug 2026)
Full report: outputs/docs/accounting-capability-research.md (current state from the code, the HMRC Self Assessment bar, gap analysis, MTD for VAT/ITSA, FreeAgent as reference, positioning recommendation, sized shortlist). It carries its own source list, with the same proxy caveat as §12 (GOV.UK unreachable directly; search-excerpt evidence cross-checked against ICAEW/ATT/ACCA/ LITRG). Headlines worth having here:
The Self Assessment bar is lower than "accounting software" implies, and we're closer than feared. Records in any format (photos fine), kept ~5 years past the 31 Jan deadline (our 6-year image retention already conservative); cash basis is HMRC's default and our stated basis; flat "equipment expensed when paid" is correct cash-basis treatment; SA103S wants nine expense boxes — or a single total-expenses figure below £90k turnover.
The structural gap is non-order income. Income derives only from Collected+priced orders — a market-stall business's walk-up takings can't be recorded at all, so the app can't hold complete income records for a big slice of the D-24 target. Shape the fix as daily takings entries — exactly the record HMRC's retail relaxations expect (VAT and MTD ITSA both).
"Letal bank" resolved: it's Mettle (as #270 already records). FreeAgent is free with NatWest/RBS/Ulster/Mettle business banking (Mettle: ≥1 transaction/month) and the free version is the full product — so competing with FreeAgent on accounting is competing with free. FreeAgent's own KB declares stock/inventory out of scope (unavailable on cash basis at all) and tells users to run a specialised system and post summaries in — we are the complement, not the competitor.
MTD for VAT: CSV export/import is explicitly a valid digital link (Notice 700/22 s4.2.1; only copy-paste breaks the chain), and bridging software is permitted indefinitely — so "export what a bridging tool needs" is a legitimate, compliant scope. Direct API filing is not worth it.
MTD for Income Tax is live (6 Apr 2026, > £50k gross; £30k Apr 2027; £20k Apr 2028, assessed on 2026-27 income — SI 2026/336). Qualifying income is gross turnover, not profit, so the 2028 wave catches real target users (~£400/week). Mandated users need commercial software even for the year-end return. The record-keeping half of the products-working-together model needs no HMRC recognition and no fraud-prevention headers — becoming "recognised software" is a permanent compliance treadmill (device-data fraud headers on every call, spec change-log, strengthening-standards programme) and not worth pursuing.
Positioning: operational record-keeper with clean handoff. For sub-threshold users, ProvenBatch alone can honestly cover Self Assessment once income is complete; for everyone else, the job is no-re-keying exports into FreeAgent/accountant/bridging software. Never claim "MTD compatible/recognised" — say "keeps the digital records; exports them for your MTD software or accountant".
Sized shortlist (full table §7 of the report): the first slice is other income/daily takings (M) + tax-year P&L (S) + accountant export pack (S–M) + SA103 box mapping on categories (S) — all one-off builds, no compliance tail. Mileage log, home-working flat rate, personal-use flag, paid-date on orders, VAT-threshold nudge are all S. Order invoice PDF is M and judged in scope (operational, not accounting). A FreeAgent-style tax set-aside estimate is the first feature with a real ongoing tail — deferred. Out of scope: bank feeds (FCA/ aggregator territory), VAT bookkeeping, any HMRC API, double-entry/balance sheet, payroll.
One dated caveat for future readers: the 2026-27 mileage flat rate rose 45p → 55p (first 10,000 miles; announced 21 May 2026, backdated to 6 Apr) — many guides still show 45p, which remains correct only for 2025-26 returns.
Full write-up
From outputs/docs/accounting-capability-research.md
Accounting capability — what "basic accounting" should cover (#341)
Researched 1 Aug 2026, against app v0.95.0. ✅ Epic #346 subsequently built the core scope (v0.116.0–v0.124.0, closed 5 Aug 2026); the VAT/export analysis remains the standing reference for what was deliberately not built. This is the full write-up for issue #341 — a research issue. RESEARCH.md §13 carries the headlines; this document carries the detail and the sizing, so a future session can turn it into scoped build issues.
The question: what is genuinely missing for a small food business to run its basic bookkeeping through ProvenBatch well enough to do a Self Assessment tax return, or hand clean records to an accountant, without keeping a second system in parallel — and how big is closing that gap?
1. What Accounts does today (v0.95.0, from the code)
The Accounts tab (index.html, renderAccounts() and around) is a monthly cash-basis money-in / money-out record:
Income is derived, not entered: orders with status = "Collected" and a price, counted in the month of collected_at (fallback due_date). collected_at is stamped when the order is advanced to Collected and is explicitly commented as the "cash-basis income date". Collected but unpriced orders raise a "heads up" strip and are excluded until priced. There is no other income path — no manual income entry, no till/stall takings, no non-order income of any kind.
Expenses are a flat categorised ledger (expenses: spent_on, category, amount, supplier_id, receipt_ref, item links). Five categories, fixed: Supplier products, Packaging, Equipment, Fees, Other (EXPENSE_CATEGORIES, CHECK-constrained in both DBs). Lines link to saved items per category (supplier products / packaging / equipment / fees), suppliers carry an optional default category (#245), and items can be moved between categories with their expenses.
Capture is strong: multi-line expense form, on-device OCR receipt scanning with AI parsing (parse-receipt), supplier auto-detection, supplier-scoped item matching, inbound email capture, and a receipt-drafts inbox. Receipt images are stored privately with the expense, deduplicated, with a retention setting defaulting to 6 years "in line with HMRC record-keeping", a review-&-purge flow and per-image "keep forever" pins.
Reporting is one month at a time: Income / Expenses / Net profit KPIs, expenses-by-category, income-by-order, searchable expense list. ‹ › month navigation only — no year view, no tax-year view (6 Apr–5 Apr), no export of the P&L, no category totals across months.
Export: the #209 full tenant export (Settings) produces a zip with one CSV per table — expenses.csv, orders.csv etc. — raw table dumps for portability, not accounting-shaped output. #337 (single multi-tab spreadsheet) and #338 (CSV import) are open follow-ons.
The page footer states the basis plainly: "Cash basis: income counts in the month an order is collected; expenses in the month you paid. A simple record for your own books — not tax advice."
What does not exist anywhere in the app: VAT (the only "VAT" in the codebase is the receipt scanner's regex for discarding summary lines), invoicing, payments/paid-status on orders (price and collected date only), bank feeds or reconciliation, chart of accounts, mileage or use-of-home records, drawings/personal-use records, capital allowances, balance sheet, tax-year anything, MTD anything.
2. What HMRC actually requires for Self Assessment (the bar to clear)
(Researched 1 Aug 2026. GOV.UK is proxy-blocked from this sandbox, so facts were assembled from search excerpts of GOV.UK/HMRC PDFs cross-checked against professional bodies (ICAEW, ATT, ACCA, LITRG). Figures worth re-verifying against GOV.UK before hard-coding are flagged.)
The bar is genuinely low — much lower than "accounting software" implies:
Records: all sales and income; all business expenses; the evidence behind them (receipts, bank statements, invoices, till rolls). No prescribed format — "on paper, digitally or as part of a software program"; photos/scans of receipts are acceptable. Records are kept in case HMRC asks (penalty up to £3,000/year for inadequate records), never submitted.
Retention: at least 5 years after the 31 January filing deadline (≈ 5¾–6¾ years after the transaction). Our receipt-image default of 6 years sits at the bottom edge of the worst case — fine in practice, and the setting already allows longer. VAT-registered businesses: 6 years.
Cash basis is now the default (since 2024-25; turnover limits abolished — any size unincorporated business). Accruals is the opt-out, ticked on the return. Under cash basis: income when received, expenses when paid, no debtors/creditors/stock valuation at year end, and equipment is simply an expense in the month paid (capital allowances survive only for cars). Our model's basis and its flat treatment of equipment are both correct, not naive.
The return itself (SA103S, turnover < £90k) wants remarkably little: turnover, and expenses in nine boxes (goods bought for resale/used; car/van/travel; staff; rent/rates/power/ insurance; repairs; professional fees; interest/bank charges; phone/stationery/office; other). Below £90k turnover there is a further easement: a single total-expenses figure is permitted with no itemisation at all. Tax years run 6 Apr–5 Apr.
Simplified expenses (optional flat rates): business mileage 45p/mile first 10,000 (2025-26), ⚠️ raised to 55p for 2026-27 (announced 21 May 2026, backdated to 6 Apr — many guides still say 45p), 25p after; working-from-home £10/£18/£26 a month by hours band (25–50/51–100/101+) — though a baker's oven-heavy electricity may beat the flat rate via actual-cost apportionment.
Trading allowance: gross trading income ≤ £1,000/year needs no registration at all; above it, register by 5 October after the tax year, then deduct EITHER the flat £1,000 OR actual expenses — never both. (A £3,000 reporting threshold was announced for "this Parliament" but is not in force — watch, don't build.)
Drawings are never an expense; goods taken for personal use under cash basis are handled by disallowing their cost (under accruals, by adding market value to income — the Sharkey v Wernher rule). Private shares of mixed costs (phone, car, home) are apportioned "just and reasonably".
VAT: registration threshold £90,000 taxable turnover in any rolling 12 months — and zero-rated food sales (most cold bread/cakes) still count toward the threshold, a genuine trap for a growing bakery. Registering makes MTD for VAT mandatory (digital records, software filing) — a step-change in bookkeeping burden.
NIC: compulsory Class 2 is gone (since April 2024); Class 4 is 6% between £12,570 and £50,270, 2% above, assessed from the return — nothing to book in-year. Below the small-profits threshold (£6,845 in 2025-26) voluntary Class 2 (~£3.50/wk) protects the State Pension — a prompt worth surfacing one day, not a bookkeeping feature.
3. Gap analysis — today's model vs "good enough" bookkeeping
Measured against §2, what we have is closer than the issue feared, with one structural hole.
Already right (and worth saying so):
Cash basis, correctly stated in the UI, is HMRC's default. Flat "equipment is an expense when paid" is the correct cash-basis treatment, not a simplification.
Expense records — dated, categorised, supplier-linked, searchable — with receipt evidence attached, privately stored, deduplicated and retention-managed already exceed what most sole traders hold. The evidence side is our strongest suit.
Income from orders is a true record of what was earned from whom, better than a bare till total.
The gaps, in order of how much they block "do a return from this app alone":
Non-order income doesn't exist. Income is only derivable from Collected+priced orders. A market-stall or farmers'-market business — squarely in D-24's target — takes most of its money as walk-up sales with no order behind them. Until takings can be recorded directly (a dated "money in" entry: stall day, till total, other income), the app cannot hold complete income records for a large slice of target users, and nothing downstream (year totals, return prep, accountant handoff) is trustworthy. This is the structural gap.
No tax-year view. The return needs 6 Apr–5 Apr totals; the app shows one month at a time with no year totals, no category totals across months, no export of either. A "Tax year" P&L (income, expenses by category, net) is the single highest-leverage addition — it converts twelve month-views into the numbers the SA103 actually asks for.
Categories don't speak SA103. Our five categories are operationally right but need a mapping to the nine SA103S boxes (Supplier products + Packaging → "goods bought for resale or goods used"; Equipment → its cash-basis expense box; Fees/Other → split across professional fees, premises, travel, other). FreeAgent's pattern (§5a) — each category carries its HMRC box, assigned once — fits perfectly, and below £90k turnover the single-figure easement means even an imperfect mapping can't produce a wrong return, only a less useful breakdown.
Expense types with no natural home: mileage (flat-rate log), home-working flat rate, and the private-use share of mixed costs. Absorbable under "Fees"/"Other" by a diligent user today, but un-prompted; a mileage log and a monthly WFH flat-rate line are small, self-contained adds.
Personal use / drawings: no way to mark goods taken for own use (cash basis: the cost should be disallowed) or record drawings (not tax-relevant on cash basis, but the number an accountant asks for first). Small.
Income date nuance:collected_at proxies "money received" — right for cash-on-collection, wrong for invoiced/paid-later customers. An optional paid-on date (or payment status) on orders would make the cash-basis date honest. Small, and pairs with any invoicing work.
No accounting-shaped export. #209 exports raw tables; an accountant wants a tax-year P&L by category, an income list, an expense list with evidence references. This "accountant pack" is the deliverable that makes "hand clean records to an accountant" literally true (see §6).
No VAT-threshold awareness. We hold rolling income; a gentle "your last 12 months' income is approaching £90,000 — VAT registration has rules worth reading" nudge is cheap insurance against the zero-rated-sales trap. ⚠️ Corrected 30 Aug 2026 (#829): the parenthetical here used to read "Actual VAT bookkeeping stays out — §4", which conflated two separable things. §4's verdict is about MTD submission and still stands. Capturing a net/VAT/gross split on an expense line is bookkeeping, not submission, and is now decided IN — outputs/docs/decisions/829-expense-capture-vat.md, tracked as #901.
4. Making Tax Digital — what it is, and whether we should be in it
(Researched 1 Aug 2026, same sourcing caveat as §2. Canonical sources: VAT Notice 700/22, the Income Tax (Digital Obligations) Regulations 2026 (SI 2026/336), HMRC Developer Hub guides — verified via search excerpts + ICAEW/ICAS/ATT/LITRG cross-checks.)
4a. MTD for VAT (mandatory for all VAT-registered businesses since April 2022)
"Functional compatible software" is defined as a set of programs joined by digital links that together keep the records digitally, compute the return from them, and submit via the VAT API. The required digital records are modest (per-supply date, net value, VAT rate; per-purchase date, value, input tax) — and retail-scheme users need only a digital record of daily gross takings, not individual sales. Two facts decide our posture:
A CSV export/import is explicitly a valid digital link. Notice 700/22 s4.2.1 lists "XML, CSV import and export, and download and upload of files" (even emailing the file) as digital links; the only named violation is copy-and-paste/re-keying. So a ledger app that exports clean CSV into a bridging tool or accountant's software is a legitimate part of functional compatible software — with no HMRC integration of its own.
Bridging software is permitted indefinitely — HMRC has stated it has no plans to stop spreadsheet+bridging, and no sunset exists in current guidance.
Verdict: our users are overwhelmingly unregistered (< £90k), and D-16 keeps even Stella Apps out of VAT. If a customer registers, the honest scope is exactly what #341 suspected: "export what a bridging tool needs" — never direct API submission. Real VAT bookkeeping (rate-coding every sale, VAT invoices, a VAT account) is FreeAgent's job.
4b. MTD for Income Tax (ITSA) — live now, and creeping toward our users
Confirmed phase-in (consolidated into SI 2026/336, in force 1 April 2026):
From
Mandatory when "qualifying income" over
Assessed on tax year
6 April 2026 (live)
£50,000
2024-25
6 April 2027
£30,000
2025-26
6 April 2028
£20,000
2026-27
Qualifying income is GROSS turnover, not profit — self-employment plus property income combined. That matters: a baker turning over ~£400/week crosses £20k, so the April 2028 wave catches a real slice of D-24's target market (and it is assessed on 2026-27 income — the records being kept now). Below £20k nothing is announced; general partnerships aren't in scope at all.
What a mandated user must do: keep digital records (each transaction: date, amount, category — categories mirroring the SA103F headings, with a retail election to record one daily gross takings figure instead of each sale, and "three-line" totals for turnover under £90k), submit cumulative quarterly updates (category totals only — no transactions — due 7 Aug / 7 Nov / 7 Feb / 7 May), then a year-end final declaration through commercial software — HMRC has confirmed its own free filing service will not be available to mandated users. Penalties are points-based (£200 at four points), with a first-year easement on quarterly-update lateness for 2026-27.
Two structural facts help us: records may live in one product and be filed through another — GOV.UK explicitly supports the "products that work together" model, spreadsheets+bridging included — and the record-keeping half requires no HMRC recognition, no Developer Hub account, and no fraud-prevention headers, because it calls no HMRC API.
4c. What becoming "HMRC recognised" actually takes
There is no HMRC fee at any stage — the cost is engineering and permanent compliance:
Developer Hub account → sandbox build (OAuth 2.0 per-user authorisation) → production approval (evidence review, up to 10 working days, six-month completion window) → listing on the GOV.UK software choices pages ("recognised", explicitly not "approved").
Fraud prevention headers are the hidden cost: a legal obligation (s135 FA 2002 directions) to collect and transmit end-user device data on every API call — for a browser SaaS like ours, Gov-Client-* headers carrying public IP/port, a minted persistent device ID, timezone, local IPs, screen geometry, browser JS user agent, plugins, MFA status, plus Gov-Vendor-* server data. That means client-side collection JS, server forwarding, UK GDPR privacy-policy wording, a validator API to pass, a change-log to track forever — and HMRC may fine or cut off vendors who persistently get it wrong.
ITSA products must meet minimum functionality standards (records + quarterly updates + signposting, or explicit pairing with another product), and HMRC published a plan in March 2026 to strengthen third-party software standards further — the treadmill is accelerating, not settling.
Verdict on recognition: not worth pursuing. It buys a listing on a GOV.UK page next to FreeAgent (free with the user's bank) and hundreds of others, at the price of a permanent compliance workstream utterly orthogonal to what makes ProvenBatch valuable. The one cost of staying out: we may never describe ProvenBatch as "MTD compatible" or "MTD recognised" — the truthful phrasing is "keeps the digital records; exports them for your MTD software or accountant".
5. FreeAgent as the reference product
(Researched via search-result synthesis 1 Aug 2026 — the sandbox proxy 403s direct page fetches, same caveat as gtm-channels-research.md. Sources are freeagent.com and support.freeagent.com KB articles via search; spot-check before quoting verbatim.)
First, the bank. #341 half-remembers "Letal bank" — it is Mettle, NatWest Group's app-based business account. Issue #270 already records it: Dave opened Mettle for Stella Apps on 29 Jul 2026 specifically because it includes the full FreeAgent platform free. FreeAgent (Edinburgh, bought by RBS in 2018 for £53m) is free with business banking from exactly NatWest, Royal Bank of Scotland, Ulster Bank and Mettle — on Mettle the condition is at least one transaction a month, and the account must have an Open Banking feed enabled (re-authorised every 90 days). Otherwise it is £19/mo ex VAT for a sole trader. Crucially the free version is the whole product — FreeAgent does not tier by feature; only add-ons (Smart Capture Unlimited £5/mo, Amazon integration) cost extra.
5a. What it does, and whether each part is relevant to ProvenBatch
FreeAgent feature
What it actually is
Relevant to ProvenBatch?
Bank feed + reconciliation
Daily Open Banking import; every transaction must be "explained" (categorised); "Guess" auto-categorises from learned behaviour + ML
Out of scope to build. Open Banking access means an FCA-regulated AISP or an aggregator contract plus 90-day reauth UX — a subsystem bigger than Accounts itself, and free in FreeAgent. The pattern (learned auto-categorisation of repeat spend) is already half-present in our supplier default categories + item matching
Partially relevant. We hold the order, the customer and the price — a printable/emailable invoice or payment-request off an order is an operational feature for our user (wedding cakes, wholesale to a café), not accounting bloat. Full invoicing machinery (recurring, credit notes, payment links) is FreeAgent's job
Expense capture
"Smart Capture": photo → extracts date + amount + suggested category only; 10/month free, then £5/mo
We already exceed it. Our scanner parses line level detail (per-item, per-supplier-product, quantities, unit prices) — FreeAgent extracts two fields and a guess. This is a genuine differentiator to keep, not a gap
VAT / MTD VAT filing
Auto-generated MTD returns filed direct to HMRC; FRS, invoice + cash basis
Out of scope (see §4)
Self Assessment filing
Auto-populates and files SA100/SA102/SA103 (short + full)/SA105 for sole traders
Out of scope to file — but the category design is the thing to borrow: every FreeAgent cost category carries a "tax reporting type" that is an SA103 box plus an allowable/disallowable flag, chosen once per category, never per transaction. Year-end totals then group themselves into the return's boxes. This is the cheapest idea in the whole product (see §7)
MTD for Income Tax
HMRC-recognised for ITSA; quarterly updates filed from web + mobile; end-of-year flows
Out of scope (see §4)
Reports
P&L, balance sheet, trial balance, aged debtors/creditors; Tax Timeline — live estimated tax owed with deadlines, exportable to calendar
P&L across a tax year is exactly what our monthly view lacks. Balance sheet / trial balance are accountancy bloat for our user. The Tax Timeline's "how much to set aside" ambient figure is repeatedly cited in reviews as the killer feature for sole traders — a borrowable idea
Payroll, dividends, CT600/Final Accounts
Full RTI payroll every plan; micro-entity accounts + CT filing
Out of scope entirely — limited-company machinery our target user doesn't have
Accountant access
Practice dashboard, per-client permission levels, accountant pulls TB/P&L/CSVs
The minimal version — "give your accountant clean exports" — is our natural ceiling (see §6)
5b. What FreeAgent deliberately does NOT do — the complement space
FreeAgent's own KB draws the line for us: its stock feature is "best suited to service-based businesses that occasionally supply small quantities of goods", is unavailable at all on cash-basis accounting (our users' default basis!), and for real inventory it tells users to run a specialised system and post monthly summary journals into FreeAgent. There is no per-product COGS, no recipes/BOM, no batches, no yields/wastage, no purchase orders, no supplier price tracking, no order management.
That is ProvenBatch's entire operational layer. The complement thesis: everything we are sits in FreeAgent's declared out-of-scope, and vice versa. The natural seam is the one FreeAgent itself names — the operational system produces periodic summaries that flow into the accounting layer. Which resolves #341's "second system in parallel" worry into a sharper question: for a user who banks with NatWest Group (a big slice of UK small businesses), the second system is free and HMRC-recognised — so the goal is not to replace it but to make ProvenBatch's records flow cleanly into it (or into an accountant's hands) with no re-keying. Competing with FreeAgent on accounting is competing with free, and losing.
6. Recommendation and positioning
Position ProvenBatch as the operational record-keeper that produces accountant-ready, MTD-import-ready output — and stay out of the filing business entirely.
The reasoning, assembled from §§2–5:
For the smallest users (under the VAT threshold, under every MTD threshold — most of the market today), Self Assessment is genuinely coverable from ProvenBatch alone: cash-basis income and expense records with evidence, a tax-year P&L, and either nine SA103S box totals or the single-figure easement. They file free via HMRC online. This is the "run your basic books here, no second system" promise we can honestly make — once the §3 gaps (chiefly non-order income) are closed.
For users with an accountant or accounting software, the second system is often free (FreeAgent via NatWest Group banking) and always better at accounting than we will ever be. Our job is the seam: clean, categorised, import-capable exports so nothing is re-keyed. Re-keying is also exactly what breaks MTD digital links — so "no re-keying" is simultaneously the convenience feature and the compliance feature.
MTD changes the timeline, not the strategy. April 2028's £20k gross threshold (assessed on 2026-27 income) will pull a real slice of target users into mandatory digital records — and the records that satisfy MTD ITSA (dated, categorised transactions; daily gross takings for retail) are the same records that close our §3 gaps. Build the ledger MTD-shaped now; never build the filing.
Do not pursue HMRC recognition (§4c). Marketing copy must say "keeps digital records and exports them for your accountant or MTD software", never "MTD compatible/recognised".
One deliberate echo of existing decisions: this is the accounting analogue of D-16 (don't VAT-register: stay under the machinery, keep the record clean) and of the labelling design rule (labels are derived from records; here, the return is derived from records — by someone else's software).
7. Shortlist and sizing
Sizing: S = a day or two inside the existing Accounts code; M = a new drawer/subsystem but established patterns; L = a genuinely new subsystem. "Ongoing" flags a maintenance or compliance tail beyond normal code ownership.
#
Feature
Size
Ongoing burden
Notes
1
Other income / daily takings — dated money-in entries alongside order income (stall day takings, one-off income), shown in the same monthly and yearly views
M
none
The structural gap (§3.1) — everything else assumes complete income. Shape it as daily takings + note — exactly the record HMRC's retail relaxations expect for both VAT and ITSA
2
Tax-year P&L — 6 Apr–5 Apr (and calendar-year) totals: income, expenses by category, net
S
none
Highest leverage per line of code; the numbers the SA103 actually asks for
3
Accountant export pack — one click: tax-year P&L, income listing, expense listing with receipt-evidence references, README; CSV (+ #337's spreadsheet when it lands)
S–M
none
Reuses #209's export machinery. This is the "hand clean records to an accountant" deliverable, and a valid MTD digital link by construction
4
SA103 box mapping on categories — each expense category carries its SA103S box (FreeAgent's pattern, §5a); export groups by box; single-figure easement note for < £90k
S
tiny (form box names)
Our five categories map today: Supplier products + Packaging → "goods bought for resale or goods used"; Equipment → cash-basis expense; Fees/Other → split. Consider whether "Fees" needs dividing (professional fees vs premises vs other)
5
Mileage log — business trips at the flat rate (45p 2025-26 / 55p 2026-27, 25p over 10k)
S
small (rates change — keep as data)
The most-claimed expense we can't record today
6
Home-working flat rate — monthly £10/£18/£26 line by hours band
S
small (rates)
Pairs with 5 as a "simplified expenses" corner; note in-app that heavy oven use may beat the flat rate via actual costs
7
Personal use & drawings — flag goods taken for own use (cash basis: cost disallowed, i.e. excluded from expense totals); optional drawings record
S
none
Correctness detail accountants ask about first
8
Paid-on date / payment status on orders — make the cash-basis income date honest for pay-later customers
S
none
collected_at remains the default; a paid date overrides
9
VAT-threshold nudge — rolling-12-month income approaching £90k → gentle warning naming the zero-rated trap
S
tiny (threshold as data)
Cheap insurance; we already hold the income series
10
Order invoice/receipt PDF — printable/emailable invoice from an order (business details, customer, items, price)
M
none
Operational, not accounting — wedding cakes and wholesale need it; label-print machinery shows the pattern. Judged in scope; payment links/recurring billing are not
11
Tax set-aside estimate — ambient "roughly £X to set aside" (income tax + Class 4) à la FreeAgent's Tax Timeline
Genuinely loved feature, but the first item with a real ongoing tail; defer until 1–4 have proven demand
Deliberately out of scope (and why): bank feeds/reconciliation (FCA-regulated open-banking access or aggregator contracts + 90-day reauth UX — bigger than Accounts itself, and free in FreeAgent); VAT bookkeeping and returns (rate-coding every sale; MTD VAT filing); any HMRC API integration (§4c's treadmill); double-entry, balance sheet, trial balance (accruals machinery our cash-basis user doesn't need); payroll; Corporation Tax; Self Assessment filing.
A sensible first build slice is 1+2+3+4: complete income, the tax-year lens, the accountant pack, and box mapping — that is the entire "Self Assessment or clean handoff" promise, all one-off builds with no compliance tail.
Sources
(All accessed 1 Aug 2026. The sandbox proxy blocks direct GOV.UK/developer-hub fetches, so GOV.UK content arrived via search excerpts cross-checked against professional bodies — same caveat as gtm-channels-research.md. Items flagged in-line above deserve a direct GOV.UK read before being hard-coded into a build.)
MTD (canonical): VAT Notice 700/22 esp. s3 (records) and s4.2.1 (digital links — CSV listed as valid, copy-paste excluded); Income Tax (Digital Obligations) Regulations 2026, SI 2026/336 (legislation.gov.uk — consolidates and revokes SI 2021/1076); gov.uk/guidance/ check-if-youre-eligible-for-making-tax-digital-for-income-tax; gov.uk/guidance/ use-making-tax-digital-for-income-tax; gov.uk/guidance/choose-the-right-software-for-making-tax- digital-for-income-tax (products-working-together model); developer.service.hmrc.gov.uk — VAT and ITSA end-to-end service guides, fraud-prevention guide + change-log, terms of use. Cross-checks: ICAEW TAXguides 01/25 & 04/25, ICAS, LITRG, Tax Adviser magazine; QuickBooks/Xero bridging pages (HMRC "no plans to stop spreadsheets" statement).
FreeAgent: freeagent.com features/pricing pages; support.freeagent.com KB — notably 360008882080 (free via NatWest/RBS/Ulster/Mettle + conditions), 115001224344 (categories carry SA103 "tax reporting types"), 115001219890 (Self Assessment scope), 5116952535954 (MTD ITSA); mettle.co.uk/ features/freeagent (one-transaction-a-month condition); FreeAgent stock-valuation KB 115001224124 (stock unavailable on cash basis; "use a specialised stock system and post journals").
Capability
Last updated 6 Sep 2026
20. SumUp integration — API capability, auth, and why it stays closed (#1107, researched 6 Sep 2026)
Full report: outputs/docs/sumup-integration-research.md (endpoint-by-endpoint capability, the OAuth scope table, sizing for two build slices, and the evidence gate). Verified against SumUp's live developer docs and their published OpenAPI spec on 6 Sep 2026, not from memory. Headlines:
Technically easier than expected. The read scope we would need, transactions.history, is a DEFAULT OAuth scope — no SumUp review, no partner agreement, no fee, and a sandbox to build against. (payments/payment_instruments need manual verification; we would need neither.) The connecting business does nothing in the SumUp dashboard — it is an ordinary OAuth consent redirect. The one real prerequisite is ours: registering a SumUp OAuth app needs "a SumUp merchant account with completed account details", i.e. Dave signing up.
The landing zone now exists, which was not true when #58 closed. #31 (orders) shipped v0.20.0 and #347 (other_income — dated money-in, built for stall takings) shipped in epic #346. So #58's "pairs naturally with #31" framing is stale, and the honest value is not #58's takings-vs-units_sold reconciliation — it is "stop making them type their takings in", one row per day into other_income.
But Dave's actual reason for closing #58 is untouched — "not clear which payment methods businesses will use… a race to integrate many." Nothing in the docs answers that; it is a customer-evidence question. The unlock is one question to the beta group (#708): which card reader do you use? A fragmented answer is itself the answer — build generic CSV import, never a provider integration.
No webhooks for card-present sales (they cover checkouts you created only), so any build must polltransactions/history with a changes_since cursor — holding a per-business, long-lived third-party credential, a class of secret ProvenBatch stores nowhere today (verified: zero refresh_token/oauth hits across every migration and all 14 edge functions). That new security surface, not the code, is what makes Slice A effort: high rather than medium. The cross-tenant polling pattern is proven (lifecycle-emails.yml), so only the credential is new.
Per-product reconciliation is possible but a trap. Line items (products[]: name, quantity, price, VAT) exist only on the detail endpoint, not the history listing — so it is N+1 per transaction, against an undocumented rate limit (429 appears exactly once in SumUp's whole OpenAPI spec, on member invitations), and products[] is populated only if the merchant rings sales through SumUp's item catalogue rather than keying in an amount. Do not build it.
Cheaper first move: generic takings CSV import (#338 is already open). SumUp exports sales history and a product-sales report with Quantity as CSV. No OAuth, no stored credential, no polling — and it works for Zettle, Square and a spreadsheet too, which answers the "race to integrate many" objection instead of sidestepping it. §13 already establishes CSV as a valid MTD digital link.
This is not a bank feed. §13 put bank feeds out of scope as FCA/aggregator territory; that reasoning does not carry over — SumUp's is a first-party merchant API authorised by the merchant to their own payment provider. Worth one confirmation before shipping, not a blocker.
Four things could not be verified from public docs and need a real account: rate limits, history retention, the /v0.1/me merchant-code endpoint (used everywhere, absent from the published spec), and whether the product-sales-report CSV is standard or POS-tier. None changes the recommendation.
Recommendation: leave #58 closed, close #1107 with the research. Revisit only on the evidence gate above — an awaiting evidence shape in WAYS-OF-WORKING §1's exact sense.
Sources (fetched 6 Sep 2026): github.com/sumup/sumup-openapi openapi.yaml · developer.sumup.com — /api, /api/transactions/list, /api/transactions/get, /api/receipts/get, /api/merchants/get, /tools/authorization/{oauth,register-app,api-keys,authorization}, /online-payments/webhooks, /webhook-docs/introduction/getting-started, /problem, /getting-started, /help · help.sumup.com — item catalogue, sales history, product sales report CSV export. Full list with what each settled: outputs/docs/sumup-integration-research.md §8.
Full write-up
From outputs/docs/sumup-integration-research.md
SumUp integration — what the API actually offers, and whether to build it (#1107)
Researched 6 Sep 2026 against app v0.331.0, verified against SumUp's live developer documentation and their published OpenAPI specification rather than from memory (CLAUDE.md #19b). This is the full write-up for #1107, a research spike — the deliverable is this document and a recommendation, not code. RESEARCH.md §20 carries the headlines; the detail and the sizing are here, so a future session can turn it into scoped build issues without re-doing the API reading.
The question: #58 ("SumUp integration… bring back payments made against a business account for validation versus sold items") was closed not_planned on 1 Aug 2026. Dave re-raised it on 2 Sep 2026 as research. So: what does SumUp's API actually let a connected business's data do for us, what would a business have to do to connect, what does it cost, where would it plug in — and is it worth building?
Recommendation, up front
Leave #58 closed and do not build this yet — but the reason has changed, and so has the size of the job.
Technically it is more feasible than expected. The read scope we would need, transactions.history, is one of SumUp's default OAuth scopes — no SumUp review, no partner agreement, no manual verification. (payments and payment_instruments need approval; we would need neither.) The connect flow is an ordinary OAuth authorization-code redirect: the business clicks "Connect SumUp", signs in, approves. No SumUp dashboard steps for the business at all.
The landing zone already exists, which was not true when #58 was closed. #31 (customer orders) shipped in v0.20.0 and #347 (other_income — dated money-in entries, explicitly for stall takings) shipped as part of epic #346. #58's stated dependency is satisfied and its framing ("pairs naturally with #31") is stale.
But Dave's own reason for closing #58 is untouched by any of this — "It's not clear which payment methods businesses will use, and this could become a race to integrate many." Nothing in SumUp's documentation answers that. It is a customer-evidence question, not a technical one, and we still have no evidence.
The build is bigger than it looks, for one specific reason: there are no webhooks for card-present sales, so this must be a polling integration holding a per-business, long-lived, revocable third-party credential — a class of secret ProvenBatch stores nowhere today (verified: zero refresh_token/oauth references across all migrations and all 14 edge functions). That is the real cost, and it is a permanent one, not a one-off build.
There is a cheaper first move that answers Dave's objection instead of ignoring it: CSV import. SumUp exports sales history and a product-sales report (with Quantity) as CSV. A generic "import a takings CSV" feature fills other_income with no OAuth, no stored credential, no polling — and works for every card provider, not just SumUp, which is precisely the "race to integrate many" that closed #58. #338 (CSV import) is already an open follow-on from the accounting work.
So: leave #58 closed with a pointer here. If the evidence arrives (§7), build Slice A only (§5) — roughly a effort: high feature slice, not a medium one. Do not build Slice B.
1. What SumUp's API actually exposes
Verified against developer.sumup.com and github.com/sumup/sumup-openapi (openapi.yaml, fetched 6 Sep 2026). The API is REST/JSON. The relevant surface for this question is small.
1.1 The endpoint that matters
GET /v2.1/merchants/{merchant_code}/transactions/history
Scopes: one oftransactions.history or transactions.read.
Query parameters that matter for reconciliation:
Parameter
What it does
changes_since
Transactions modified at or after an ISO8601 timestamp — the incremental cursor
oldest_time / newest_time
Created at-or-after / before a timestamp — the date-window filter
oldest_ref / newest_ref
Event-reference cursors; supersede the time parameters when both are given
Response is { items: [...], links: [...] } — links carries the pagination cursors.
Two things are worth pulling out:
CASH is a payment type. SumUp's app lets a merchant record cash sales alongside card ones, so a business that rings everything through SumUp would give us a genuinely complete daily takings figure, not just the card half. That is better than it first appears for the tax-records use case (§4).
changes_since exists, which means an incremental poll is cheap and correct — including picking up refunds and chargebacks against sales we already imported.
1.2 What a history item contains — and what it does not
Each item in the history listing is TransactionHistory: transaction_id, transaction_code, amount, currency, timestamp, status, payment_type, card_type, type, user (the merchant email), installments_count, payout_date, payout_type, refunded_amount, plus product_summary — a free-text one-liner taken from the related checkout's description.
🔴 The history listing carries no line items. There is no products array on TransactionHistory. It gives you money and time, not what was sold.
1.3 Line-item detail exists, but on a different endpoint
GET /v2.1/merchants/{merchant_code}/transactions?id=… | transaction_code=…
The detailed TransactionFull response does carry a products array, described in the spec as "List of products from the merchant's catalogue for which the transaction serves as a payment", with per-line name, description, price, quantity, vat_rate, vat_amount, total_price, total_with_vat. There is a vat_rates breakdown and a transaction_events array (payouts, refunds) alongside it. GET /v1.1/receipts/{transaction_id} exposes the same product detail.
This is the finding that decides the shape of any build, and it cuts both ways:
✅ A real per-product units_sold cross-check is possible. #58's original ask is not fantasy.
❌ It costs one extra API call per transaction — the list gives totals, the detail gives lines. A market stall doing 80 sales a day is 80 extra calls, per business, per day, against an undocumented rate limit (§3).
❌ products is only populated if the merchant rings the sale through SumUp's item catalogue. A business that keys in "£4.50" and taps charge — which is how a great many small food businesses actually use a card reader — produces a transaction with an amount and no products at all. We would be building a reconciliation feature whose core data is optional and invisible until after a business connects.
1.4 Everything else on the API
Payouts (GET /v1.0/merchants/{merchant_code}/payouts) — what SumUp actually settled to the bank, net of fees. Useful for bank reconciliation; not what #58 asked for.
Checkouts, Readers, Customers, Members, Roles — for taking payments and running a till. Out of scope: ProvenBatch is not becoming a payment terminal.
⚠️ There is no catalogue endpoint in the published OpenAPI spec. The products scope exists ("View and modify products, shelves, prices, and VAT rates") but no /products path is in sumup-openapi. So we cannot read a merchant's item list up front to build a mapping screen — we would have to discover their product names by observing transactions.
⚠️ Merchant-code discovery is not in the published spec either. Every transactions call needs {merchant_code}. GET /v0.1/me is widely used to obtain it and appears in SumUp's own SDK issue tracker, but it is not a documented path in sumup-openapi. Building on it means building on an undocumented endpoint. This needs confirming against a real account before any build.
1.5 Webhooks do not cover this
SumUp's webhooks notify on status changes for a checkout you created, subscribed via a return_url on checkout creation, signed HMAC-SHA256. Real-time transaction results are available "via webhooks with the Cloud API" — that is the reader-integration product, for applications driving a terminal.
🔴 There is no webhook for "this merchant took a card payment on their own reader." Any ProvenBatch integration must polltransactions/history with a changes_since cursor per connected business. This is the single biggest driver of the ongoing cost in §5.
2. The auth model — what a business would actually have to do
2.1 Two models, and only one fits
Model
Fit
API keys — static credentials created by the merchant at me.sumup.com → Settings → For Developers → Toolkit → API Keys
❌ Wrong for us. They "grant broad access to the merchant account", and SumUp says explicitly "Do not expose secret API keys in publicly accessible places". Asking a customer to paste a full-access key into ProvenBatch is a bad ask and a worse thing to store. Fine for a one-off Dave-only spike; not a product.
OAuth 2.0 authorization code — "Standards-based authorization for multi-merchant solutions. Use it when other merchants or their employees connect to your application and must explicitly grant access."
✅ Exactly our case.
2.2 The OAuth flow, verified
Authorize: https://api.sumup.com/authorize — response_type=code, client_id, redirect_uri, scope, state
Access token lifetime: 3599 seconds (~1 hour). Refresh tokens are long-lived; the refresh response may return a new refresh token, which must be persisted. An invalid_grant on refresh means the token is dead and the whole authorization flow must restart.
2.3 The scopes, and the good news
Scope
Default?
What it grants
transactions.history
Yes
"View transactions and transaction history for the merchant user"
user.app-settings
Yes
SumUp mobile app settings
user.profile_readonly
Yes
"View profile details of the merchant user"
user.profile
No
Modify profile
user.subaccounts
No
Sub-account profiles
user.payout-settings
No
Payout settings
products
No
Products, shelves, prices, VAT rates
payments
No
Requires manual verification by SumUp
payment_instruments
No
Requires manual verification by SumUp
🟢 We would need transactions.history and user.profile_readonly. Both are default scopes. No SumUp approval process, no partner agreement, no commercial negotiation. This is a materially lower barrier than most financial-data integrations and is the strongest single point in favour of building.
2.4 What the business actually does — plainly, per CLAUDE.md #1/#19b
The connecting business does nothing in the SumUp dashboard. Their whole journey is:
In ProvenBatch: Settings → Connect SumUp.
They are redirected to SumUp and sign in with their normal SumUp credentials.
SumUp shows its own consent screen naming ProvenBatch and the permissions requested.
They approve; SumUp redirects back to ProvenBatch; we exchange the code for tokens.
They can revoke at any time from their SumUp account, and if they do, our next refresh returns invalid_grant and we must handle it gracefully rather than silently stopping.
2.5 What we would have to do first
ProvenBatch would have to register as a SumUp OAuth application, at me.sumup.com/settings/oauth2-applications. SumUp's documented prerequisites, verbatim:
"Prepare a SumUp merchant account with completed account details."
"Prepare one or more redirect URIs. SumUp redirects users to these URIs after authentication…"
Required at registration: application name and homepage, client name, application type (Web/Android/iOS/Other), and redirect URI(s). Optional: logo, terms URL, privacy policy URL, authorized JavaScript origin. Multiple client secrets can be created per app.
⚠️ This is a real prerequisite with Dave's name on it. It requires a SumUp merchant account with completed account details — i.e. Dave signing up to SumUp as a business. It is a dashboard step, it is genuinely his (an account signup), and it is the one thing in this whole integration that cannot be automated. Registering the redirect URIs is also environment-sensitive: staging and production have different hostnames (staging.provenbatch.co.uk / app.provenbatch.co.uk) and both would need registering, or we would need two apps.
No pre-approval, partner agreement or verified status is documented beyond "completed account details".
3. Pricing and rate limits on the API itself
Pricing: no fee is published for API access. SumUp's developer portal, FAQ and getting-started pages carry no developer, integration or per-call charge. The only money in the relationship is SumUp's ordinary card-processing fee, which the business already pays on every sale and which is out of scope for this question. Sandbox merchant accounts exist, mirror real accounts, and process no real funds — so building and testing costs nothing.
🔴 Rate limits: undocumented, and I could not verify a number. This is a real finding, not a gap in the reading:
There is no rate-limits page on developer.sumup.com.
SumUp's own Problems reference lists 20 error types and none of them is a rate-limit / 429 / too-many-requests type.
In the published OpenAPI spec, 429 appears exactly once across the whole file — on the member invitation endpoint (POST …/members), with a Retry-After header. It appears nowhere on the transactions endpoints.
The honest reading is that a limit almost certainly exists and is simply not published. Anything we build must therefore assume an unknown ceiling: conservative polling, exponential backoff, Retry-After respected wherever it appears, and no design that scales calls linearly with transaction volume if it can be avoided. That last point is exactly what makes §1.3's per-line detail fetch expensive — it is the one design choice that scales with how busy a business is, which is the worst possible thing to scale on against an unknown limit.
Data retention — how far back transactions/history reaches — is also undocumented and could not be verified. It matters for the first-connect backfill and needs a real account to answer.
4. Where this would plug into ProvenBatch
The picture has changed materially since #58 was triaged, and this is the part of #58 that is genuinely stale.
4.1 What already exists (from the code, v0.331.0)
batches.units_sold — an integer the user types by hand in the batch detail drawer (bd_sold), alongside sale_price_each. It feeds the unsold/waste figures (Math.max(0, b.qty - b.units_sold)) and the made-vs-sold analytics, and only counts where both qty and units_sold are recorded. This is #58's "sold items".
Orders (#31, shipped v0.20.0) — orders / order_items / customers, with per-line unit_price since #1063. Income is derived from priced orders on a cash basis via incomeDateOf() = paid_on (#348) → collected_at → due_date.
other_income (#347) — dated money-in entries independent of orders: received_on, amount, description. Built explicitly for "stall takings, one-off income". It flows into the Accounts monthly summary and the accountant export.
4.2 The gap SumUp would fill
outputs/docs/accounting-capability-research.md names the structural gap precisely: income was derived only from collected priced orders, so "a market-stall business's walk-up takings can't be recorded at all". #347 closed that gap manually — someone types the day's takings in.
SumUp's honest value is not #58's framing. It is: stop making them type it.
That reframes the integration as "automatically fill other_income from the card reader" — a one-row-per-day insert from a single transactions/history call, which is a far smaller and far more defensible feature than "reconcile takings against units_sold".
4.3 Is the #58 framing still right? No.
#58 said this "pairs naturally with #31 customer orders — on its own it validates takings against units_sold." Three problems with that today:
#31 shipped. The dependency is satisfied, so it no longer defers anything.
Orders and SumUp barely overlap. Orders are pre-orders with a fulfilment flow (New → Confirmed → Ready → Collected). A SumUp transaction is a counter sale. An order collected and paid by card would appear in both, which makes naive import a double-counting risk against order_income, not a synergy. Any build needs a deliberate answer to "is this SumUp transaction already recorded as an order?" — and the API gives us nothing to match on but amount and timestamp.
**units_sold is per batch, SumUp's quantities are per product.** A batch is one making of one recipe; a product is the sellable good. Reconciling the two needs a batch↔product↔SumUp-item mapping, and the SumUp half of that mapping is free text typed by the business into a different app.
**So the standalone "show takings vs. units_sold" view is the harder of the two shapes, not the simpler one.** The simpler, more valuable shape is takings → other_income.
4.4 One thing genuinely in our favour
This is not a bank feed.accounting-capability-research.md put bank feeds out of scope as "FCA/aggregator territory", and that reasoning does not carry over: SumUp's API is a first-party merchant API, authorised by the merchant to their own payment provider, not account information services over someone else's bank account. That is a materially different regulatory shape. (Noted as a distinction worth relying on, not as legal advice — worth one confirmation before shipping.)
4.5 And one thing genuinely against
SumUp already does a lot of this. Its item catalogue supports categories, variants, prices, taxes, barcodes and inventory tracking, syncing across Terminal, POS Lite, Register, Kiosk, Online Store and the Business app. A merchant who rings items through SumUp can already see units sold, in SumUp, without us. The ProvenBatch-specific value is narrower than #58 implies: it is closing the made → sold → wasted loop against batches and recipe cost, which SumUp cannot do. That is a real value, but it only exists for a business already disciplined enough to keep the SumUp catalogue accurate — and a business that disciplined is exactly the one already getting half the answer elsewhere.
5. What it would cost to build
Two slices, and they are very different sizes. The temptation is to describe one feature; they are not one feature.
Slice A — takings import (totals only)
"Connect SumUp, and your daily card and cash takings appear in Accounts."
Piece
Notes
OAuth app registration
Dave's SumUp signup (§2.5) + redirect URIs for staging and production
Connect/disconnect UI
A Settings card: connect, show connection status, disconnect. Two states plus an error state ("SumUp access was revoked — reconnect")
Per-tenant token storage
New table, RLS-scoped, holding refresh_token + merchant_code. This is the new thing — see below
Token refresh
~1-hour access tokens; handle rotated refresh tokens; handle invalid_grant by marking the connection dead and telling the user
Daily poll
Scheduled sweep across connected businesses on both projects, changes_since cursor per business, backoff against an undocumented limit
Import + dedupe
One other_income row per business per day; idempotent on re-run; handle later refunds/chargebacks amending a day already imported
Double-count guard
Don't silently inflate income when a collected order was also paid by card (§4.3)
Tests + UAT
Suite coverage, and the harness for the Settings flow, mobile and desktop
Two things push this above a normal effort: medium:
🔴 A new class of stored secret. Verified: ProvenBatch stores no per-tenant third-party credential anywhere today — zero refresh_token/oauth hits across every migration and all 14 edge functions. Stripe is platform-level (our own account), not per-tenant. A SumUp refresh token is a long-lived, per-business credential granting read access to that business's money. It needs encryption at rest, RLS that is right first time, a revocation path, an entry in outputs/docs/credential-inventory.md, and a genuine answer to "what happens when it leaks". This is a new security surface for the product, and it does not go away after the build.
🟢 The polling harness pattern already exists. Cross-tenant scheduled sweeps across both Supabase projects using service-role keys are proven — lifecycle-emails.yml, feedback-status-sync.yml and ~28 registered ops jobs do exactly this shape. So the mechanism is not new; only the credential is.
Size: effort: high. Not because any one piece is hard, but because it is a feature slice plus a new schema plus a new ongoing job plus a new security surface, and the last of those carries a permanent maintenance tail.
Slice B — per-product units_sold reconciliation
Everything in Slice A, plus:
A detail fetch per transaction (§1.3) — N+1 against an undocumented rate limit
A SumUp-item ↔ ProvenBatch-product mapping screen, built from names observed in transactions because there is no catalogue endpoint (§1.4)
Ongoing mapping maintenance every time the business renames or adds a SumUp item
A reconciliation view, and a considered answer to what a discrepancy means (theft? waste? unmapped item? cash sale not rung through? a sale from a batch made three days ago?)
Batch↔product attribution: which batch does a sale on Tuesday draw down? Nothing in the data says.
Size: genuinely large, and the payoff only exists for the subset of merchants who ring every sale through a well-maintained SumUp catalogue. The discrepancy question is the killer: a number that does not reconcile, with no way to say why, is a support burden rather than a feature.
Do not build Slice B. If Slice A ships and businesses ask for product-level detail, revisit it with real data about how they actually use their reader.
Slice 0 — the cheaper thing that may be better than either
Generic takings CSV import. SumUp exports sales history from the dashboard, and its product sales report exports a CSV including Quantity, sales incl/excl tax, and tax amount. (One caveat to confirm: the richer product-sales-report export is documented under SumUp's POS product pages, so it may be a paid-tier feature rather than something every SumUp user has. The basic sales history export is standard.)
A "import a takings CSV into other_income" feature has: no OAuth, no stored credential, no polling, no token refresh, no revocation handling, no new security surface — and it works for Zettle, Square, Stripe Terminal and a spreadsheet too. That is a direct answer to the exact objection that closed #58, rather than a way around it. #338 (CSV import) is already an open follow-on from the accounting epic, and accounting-capability-research.md establishes that CSV is an accepted digital link for MTD purposes.
Size: medium. It is the highest value-per-unit-of-risk option on this page.
6. The recommendation
Leave #58 closed. Close #1107 with this document. Do not build a SumUp API integration now.
The reasoning, in order:
The blocker is evidence, not capability. Everything technical checks out better than expected — default scopes, no approval, no fee, a sandbox, and a landing zone (other_income) already built. What we still cannot answer is Dave's own question from 1 Aug: do our businesses actually use SumUp? Building a provider-specific integration on a guess is how the "race to integrate many" starts.
The honest value is smaller than #58 implies. The per-product reconciliation #58 describes is optional data (§1.3), needs a mapping screen we cannot pre-populate (§1.4), and produces discrepancy numbers we cannot explain (§5B). The defensible value is "stop making them type their takings in" — real, but it is an ergonomics win on a feature that already works manually.
The cost has a permanent tail. Polling infrastructure is fine; we do that already. A per-tenant third-party credential is a new security surface for a compliance product, and it stays new forever.
A better first move exists. CSV import gets most of the value, for every provider, at lower risk, and answers the objection instead of sidestepping it.
This is a case for awaiting evidence — WAYS-OF-WORKING §1's exact definition: not blocked, not forgotten, needs real customer signal first, and the issue names the evidence that would unlock it.
7. What would change this answer
Named concretely, so a future session does not have to re-derive the judgement:
🔓 The unlock: ask the beta group what they take payments on. One question to the tester community (#708) — "which card reader or payment app do you use?" If SumUp is clearly dominant among respondents rather than one of five answers, Slice A becomes worth building and the "race to integrate many" objection is answered by data. If the answer is fragmented, build Slice 0 (CSV import) and never build a provider integration at all — the fragmentation is the answer.
A business asking for it by name, in-app feedback or otherwise, is worth more than any of this analysis.
SumUp publishing rate limits and history retention would remove two unknowns that currently force conservative design (§3).
SumUp publishing a catalogue endpoint would make Slice B's mapping screen tractable. Today it is not.
Two things that should not move this: SumUp's API being easy (it is, and that is not a reason to build something), and the integration being "only" a few days' work (Slice A is not, and the tail is permanent).
8. Verification notes and sources
Everything in §1–§3 was read from SumUp's live documentation on 6 Sep 2026, not from model memory, per CLAUDE.md #19b. Where a fact could not be verified it is marked as unverified above rather than filled in — specifically: rate limits, transaction-history retention, the /v0.1/me merchant-code endpoint, and whether the product-sales-report CSV is standard or POS-tier. All four need a real SumUp account to settle and none of them changes the recommendation.
Primary sources:
github.com/sumup/sumup-openapi — openapi.yaml (fetched 6 Sep 2026): endpoint paths, the TransactionHistory / TransactionFull / Product / PaymentType / EntryMode schemas, the 429-appears-once finding, and the absence of any /products or /v0.1/me path.
developer.sumup.com/api — API reference index (Checkouts, Readers, Customers, Transactions, Payouts, Receipts, Members, Memberships, Roles, Merchants).
developer.sumup.com/api/transactions/list and /api/transactions/get — the history and detail endpoints, scopes, filters, and the products[] array.
developer.sumup.com/api/merchants/get — merchant profile fields and scopes.
developer.sumup.com/tools/authorization/oauth — endpoints, grants, the 3599-second token lifetime, refresh behaviour, and the full scope table with defaults.
developer.sumup.com/tools/authorization/register-app — registration prerequisites (verbatim), required fields, and which scopes need manual verification.
developer.sumup.com/tools/authorization/api-keys — the API-key model and SumUp's own security warnings.
developer.sumup.com/tools/authorization/authorization — the three authorization models.
developer.sumup.com/online-payments/webhooks and /webhook-docs/introduction/getting-started — webhook scope (checkout status only) and HMAC-SHA256 signing.
developer.sumup.com/problem — the 20 documented problem types, none rate-limit related.
developer.sumup.com/getting-started and /help — sandbox merchant accounts; no published API fee.
help.sumup.com — item catalogue (categories, variants, prices, taxes, inventory tracking, sync across SumUp surfaces); sales history and product-sales-report CSV exports and their columns.
ProvenBatch sources (read from the tree at v0.331.0):outputs/bakery-app-site/src/part-2.js (units_sold / bd_sold, incomeDateOf, otherIncomeInRange, other_income table mapping); outputs/migrations/347_other_income_*.sql, 348_orders_paid_on_*.sql, 1063_order_line_prices_*.sql, 479a_batch_products_sessions_*.sql; outputs/docs/accounting-capability-research.md; outputs/RESEARCH.md §13; issues #58, #31, #347, #348, #1063; .github/workflows/lifecycle-emails.yml and outputs/ops/jobs.d/ for the existing cross-tenant sweep pattern; and a repo-wide search for refresh_token/oauth across outputs/migrations/*.sql and supabase/functions/*/index.ts (zero hits) for the "new class of stored secret" finding.
26. How much of Claude and Anthropic could move to xAI (researched 13 Sep 2026)
The finding: do not do a wholesale Claude → Grok cutover. ProvenBatch uses Anthropic in three different businesses, and they transfer at very different rates. Full write-up: outputs/docs/claude-to-grok-transfer-research.md.
Stack
Today
To xAI / Grok
A. In-app product AI
Anthropic Messages in edge functions
Mostly transferable, with one hard quality gate (parse-pack)
B. Dev / shipping
Claude Code + Grok + Codex since #1563
Already ~80% transferred. Grok is a full orchestrator
C. Ops routines
Claude Routines + MCP connectors
Partially transferable, connector-limited. beta-updates-digest already runs as a Grok automation
Quality: pack-photo transcription (parse-pack, Opus vision) must not move without a gold-set eval — a fluent hallucination of a hidden allergen is a labelling incident. ProvenBot compliance and receipt/recipe OCR need the same kind of eval at lower stakes.
Capability we would lose without extra work: inline PDF document blocks, webp/gif images, Claude.ai cloud sessions, Gmail- and Metricool-connected routines, Anthropic’s 90% prompt-cache discount on ProvenBot. Adding xAI to product paths is a new sub-processor (clause 3.2 notice before go-live).
If proceeding: gold-set eval before touching parse-pack or ProvenBot compliance; optional cheap xAI wedge on recall-watch explain; leave Claude Code in the mix; port ops routines one connector family at a time (not via scheduler_create, which expires in 7 days). Reopens #1189 and the parse-pack “do not shrink the model” rule.
Corrected 1 Oct 2026 (#2988): the original pass also scored a second inference vendor for voice notes and cheap vision. Voice left the product (#1939) and that account is closed, so those columns are removed here and in the full write-up.
From outputs/docs/claude-to-grok-transfer-research.md
Claude and Anthropic — what transfers to xAI / Grok
Researched 13 Sep 2026. Companion to RESEARCH.md §26. Not a decision and not a runbook — nothing has been switched.
The question: how much of ProvenBatch's Anthropic / Claude capabilities could move to xAI / Grok, and whether quality or capability would be lost.
Corrected 1 Oct 2026 (#2988): the original pass also scored a second inference vendor (speech-to-text for voice notes, and cheap vision). Voice left the product in #1939 and that account is now closed, so its columns, its section and its recommendations are removed. The xAI findings are unchanged.
0. The answer in one page
Do not do a wholesale Claude → Grok cutover. ProvenBatch uses Anthropic in three different businesses, and they transfer at very different rates:
Stack
What it is today
Transfer to xAI / Grok
A. In-app product AI
Anthropic Messages API in edge functions
Mostly transferable, with one hard quality gate
B. Dev / shipping (Claude Code)
Already shared with Grok and Codex since #1563
Already ~80% transferred. Grok is a full orchestrator. The remaining gap is cloud sessions
C. Ops routines
Claude Routines + MCP connectors (Gmail, Metricool, Calendar, Drive, Cloudflare)
Partially transferable, connector-limited. Dave already declined new Claude Routines and put beta-updates-digest on a Grok automation
Quality loss is real in one place and speculative in two. Pack-photo transcription (parse-pack, Opus vision, allergen incident if wrong) must not move without a gold-set eval on real packs. ProvenBot compliance answers and receipt/recipe OCR need the same kind of eval, at lower stakes. Everything else is a rewrite plus a legal notice, not a quality cliff.
Historical context that matters. Ops originally ran on Grok Bot (roster exported 25 Aug 2026). #709 moved those jobs onto Claude Routines on purpose — task-centred, repo-versioned skills, connectors already wired. Asking “can we move back to Grok” is a reversal of a completed migration, not a greenfield. The product AI (receipts, packs, ProvenBot) never ran on Grok.
1. What ProvenBatch actually uses today
1a. Product (customer data, billed per call)
All of these are Supabase edge functions. Keys live as function secrets, never in the client. Caps and spend ceilings are in-function (ai_limits). Every Anthropic call is https://api.anthropic.com/v1/messages.
Surface
Model
Why this model
Anthropic-specific API
Fallback if key missing
parse-receipt
claude-haiku-4-5 vision
Cheap; a wrong line costs pennies
Image and PDF document blocks; jpeg/png/webp/gif
No on-device reader — failed draft, type by hand
parse-recipe
claude-haiku-4-5
Same
Photo, PDF document, or pasted text
On-device parser, and it says so
parse-pack
claude-opus-5 vision
“A declaration read wrong is an allergen incident.” Volume is low (once per supplier product). README forbids shrinking the model
Multi-image, jpeg/png/webp/gif; verbatim JSON; refuse if illegible
On-device OCR exists but was unreachable on outage until #1299
Keyword test inside the function, never from the browser
Streaming SSE; system block with cache_control: ephemeral on a ~16k-token FSA corpus (cached read = 1/10 input — “the entire cost model”); client-executed read-only tools
Stops answering
admin-api drafts
claude-opus-5
Layout fabricates; brand-product transcribes verbatim like parse-pack
Vision on some drafts
Plain error, no crash
recall-watch.yml
Anthropic optional
Explain a match; matching itself is deterministic (recall-match.js)
Messages API
Falls back to plain-text matching
Also: ANTHROPIC_ADMIN_KEY drives anthropic-cost.yml and anthropic-key-audit.yml (nightly spend snapshot + workspace/expiry audit). That is ops observability, not customer-facing.
1b. Dev tooling (Claude Code / Grok TUI / Codex)
Settled 9 Sep 2026 (#1563): all three tools share AGENTS.md. The capability matrix in WAYS-OF-WORKING.md §3.2a:
Invariant
Claude cloud
Claude local
Grok
Codex
Durable named work
Yes — create_session
No — no claude-code-remote
Yes — spawn_subagent
No
Stall detection
Yes — get_session
No
Partial — subagent status only
No
Wake / queue holder
Yes — send_later / Routines
No
Yes — scheduler_create
No
Issue / PR write
Yes
Yes — gh
Yes — GitHub MCP
No
Workflow dispatch
Yes
Yes — gh
Yes — actions_run_trigger
No
Grok already qualifies as orchestrator. Local Claude does not. Codex is a worker whose GitHub half someone else carries.
What is still Claude-Code-specific and load-bearing (AGENTS.md routing table): .claude/ is tracked — settings.json permissions + hooks, production-DDL guard, session/branch-naming reminder, /batch, ops-routine skills. Grok has the analogue: .grok/skills/batch/SKILL.md, SessionStart stale-rules hook (#1618), GitHub MCP, scheduler_create.
1c. Ops routines (Claude.ai, not the API)
Judgement jobs run as Claude Routines created in the claude.ai UI so they get MCP connectors. A session-created trigger cannot be given connectors (outputs/ops/routine-setup.md, proven 26 Aug). Live routine jobs: daily-ops-brief, work-distribution, roadmap-refresh, support-triage (Gmail + GitHub), weekly-review, release-verification, connector-health, outreach-upkeep, metricool-draft-scheduling (Metricool connector), usage-review, competitor-research-refresh.
Already moved off Claude:beta-updates-digest is a Grok automation (82f0b568-54fa-4577-bcd9-e2e68b12a501). Notes say Dave declined new Claude Routines after the cutover.
#1189 (6 Sep) decided “ops routines stay Claude Routines.” That decision and the later “no new Claude Routines” note are in tension. A full transfer would reopen #1189.
2. What xAI / Grok can actually do (docs, 13 Sep 2026)
Inference is OpenAI-compatible at https://api.x.ai/v1. Default model grok-4.6.
Need in ProvenBatch
xAI today
Fit
Text + tools + streaming
Responses / chat completions, function calling, SSE
Can replaceANTHROPIC_ADMIN_KEY jobs, with a rewrite. Better scoping than Anthropic’s unscoped admin key
Web search
Built-in web_search tool, $5 / 1k calls
Better native fit than Anthropic-plus-search
UK GDPR transfer
xAI DPA (effective 9 Jun 2025) incorporates SCCs + UK Addendum. Processing is US; Voice docs claim GDPR / SOC 2 Type II / optional EU residency. eu-west-1 exists on grok-4
Same class as Anthropic (US + UK Addendum). Adding xAI to product paths is a new sub-processor: public page + subscriber notice before it processes (clause 3.2)
Not equivalent, and do not paper over it:
Image types: client currently accepts webp/gif. xAI would 400 those unless we transcode.
No Anthropic-style PDF document block.
Prompt cache is prefix-automatic, not breakpoint-explicit. ProvenBot’s corpus-as-system-block still works if it stays the stable prefix.
scheduler_create in Grok TUI expires in 7 days and is session-scoped. It is not a Claude Routine. Persistent Grok automations exist (the digest uses one) but are a different product surface, with their own connector story.
Grok TUI MCP is whatever is in ~/.grok/config.toml. GitHub is wired. It does not automatically inherit Claude.ai’s Gmail / Metricool / Calendar / Drive / Cloudflare connectors.
3. Transfer score, surface by surface
Keep = do not move. Move = rewrite is mechanical, quality likely comparable. Eval-then-move = API can do it, quality is the question. Cannot = missing capability or legal/ops blocker.
Stack A — in-app product
Surface
→ xAI Grok
Quality / capability loss
parse-pack (Opus, allergens)
Eval-then-move, and only onto grok-4.6 (or better)
Highest risk in the repo. Prompt is “verbatim or null, never reconstruct.” A model that “helpfully” completes a glare-hidden line is a labelling incident. No published OCR-on-curved-UK-pack gold set for Grok. Do not ship on a blog-benchmark.
parse-receipt
Eval-then-move (grok-4.3 or 4.6)
Medium. Wrong line is caught at reconciliation. Lose webp/gif unless transcoded. Lose inline PDF unless Files API or rasterise-first.
parse-recipe
Same as receipt, plus on-device fallback already exists
Medium-low. Fallback hides a bad model more than receipts do.
provenbot routine
Move onto grok-4.3 or grok-build-0.1
Streaming + tools transfer. Cache discount shrinks (Sonnet 0.1× vs grok-4.6 0.25×). Use grok-4.3 to keep cost in the same band.
provenbot compliance
Eval-then-move onto grok-4.6 only
Legal Q&A over a curated FSA corpus. Design doc forbids “am I compliant.” Still: a worse model will sound confident. Needs a frozen question set with citations.
admin-api drafts
Move
Brand-product draft is the same verbatim rule as parse-pack — inherit that eval.
recall-watch explain
Move
Optional, fail-soft.
Cost / key audit
Move (Management API)
xAI has a usage API. Rewrite anthropic-cost.yml / anthropic-key-audit.yml.
Sub-processor / DPA
New xAI row + notice before go-live
Clause 3.2 is a product commitment. Solicitor review is still deferred (#211).
Price sketch (not a reason to move pack reads): grok-4.6 is $2 / $6 per 1M vs Opus $5 / $25 — cheaper if quality holds. Haiku $1 / $5 vs grok-build-0.1 $1 / $2. ProvenBot’s cached corpus is the one place Grok 4.6 is slightly dearer than Sonnet-with-cache unless we drop to grok-4.3.
Stack B — coding / shipping
Capability
Transfer
Loss?
Batch orchestrator
Already on Grok (SKILL.md, all five invariants)
Stall detection is weaker (get_session vs subagent status). Wake is scheduler_create not send_later. Protocol is the same.
Worker quality
Grok 4.6 as worker is unmeasured against the Opus-spawn / Sonnet-stall record (#1229: 2/7 Sonnet stalls)
Unknown, not proven worse. Do not assume Grok inherits Claude’s “spawn on Opus” lesson.
Local checkout
Grok is better than local Claude (local Claude has no claude-code-remote)
Skill is .agents/skills/impeccable/, symlinked for Claude
Portable.
GitHub Claude app / claude-code-action
N/A — #1092 declined putting Anthropic keys in Actions for PR review
Nothing to transfer.
Net: stack B is already the hybrid Dave asked for in #1563. Killing Claude Code would lose cloud sessions, not the batch protocol.
Stack C — ops routines
Job family
Connectors it needs
Grok path
Loss
Daily brief / work distribution / roadmap
GitHub, Calendar, Supabase
GitHub MCP + gh + curl: yes. Calendar: only if a Google MCP is added
Push+email completion notifications are Claude Routine product. Grok automations deliver in chat (the digest already does).
Support triage
Gmail + GitHub
Needs a Gmail MCP in Grok. Not present today
Capability loss until Gmail is wired. This is the one routine whose Gmail path is proven live.
Metricool drafts
Metricool
Needs that connector on Grok
Capability loss until wired.
Outreach
Gmail
Same as support
Same.
Connector-health
Claude Code Remote + the rest
Circular: it exists to watch Claude connectors
Rewrite around Grok MCP health, or keep one Claude routine as the watchdog.
Persistence
claude.ai Routines UI, durable
scheduler_create is 7-day / session. Grok automations (digest) are the real analogue
Using the TUI scheduler as a Routine replacement would silently die. Must use automations, not scheduler_create.
Net: skills (markdown) transfer in an afternoon. Connectors do not. Support triage and Metricool are the two that actually fail closed without Claude.ai.
4. Would we lose quality?
Where we almost certainly would, unless an eval says otherwise
parse-pack. The repo’s own words: do not optimise this to a smaller model. Grok 4.6 is frontier-class and might match Opus on verbatim OCR of small, curved, glared UK packs — that is a test, not a belief. Failure mode that matters: fluent hallucination of a hidden allergen, which the prompt tries to forbid with legible: false. A model that hates saying “I can’t read this” is worse than a dumber model that refuses.
ProvenBot compliance tier. Same pattern: sounds-right legal answers with a citation that isn’t quite the FSA page. Design already forbids “am I compliant”; it does not forbid a wrong explanation of the 14 allergens.
Where we probably would not
Receipt lines, recipe import, recall explanations, admin layout drafts — Haiku-class or draft-and-human-paste. Grok 4.6 should be at least as good as Haiku.
Batch orchestration mechanics — already running on Grok.
Where we would lose capability, not quality
Inline PDF document blocks (receipts + recipes) unless rewritten to xAI Files or rasterisation.
webp/gif pack/receipt photos unless transcoded to jpeg/png.
Claude.ai cloud sessions and artifacts.
Gmail- and Metricool-connected routines, until those MCPs exist on Grok.
Anthropic prompt-cache 90% discount (ProvenBot cost model), unless we pick a cheaper Grok SKU.
Where we would gain
Native STT/TTS/Imagine if we ever want them in-product (new sub-processor uses).
Native web search.
Scoped API keys and a real usage API on xAI (Anthropic’s admin key is unscoped).
Local orchestration that Claude local cannot do.
5. Suggested path if we proceed (not this sitting)
A wholesale cutover is the expensive way to learn pack-OCR is worse. Sequence:
Do not touch parse-pack or ProvenBot compliance until a gold set exists. ~30 real pack photos (glare, folds, own-brand, multi-panel) plus the frozen ProvenBot question set with expected citations. Score verbatim match / legible:false honesty / allergen emphasis, not “looks plausible.”
Optionally A/B Haiku vs grok-4.3 on parse-receipt only (fail-soft, pennies).
Optional cheap xAI wedge:recall-watch explain. New XAI_API_KEY function secret, inventory entry, sub-processor notice before any customer content (receipts, packs, chat) hits xAI.
Leave Claude Code in the mix. Grok already orchestrates. Retire Claude only if cloud sessions are explicitly surplus.
Ops: port remaining routines to Grok automations one connector family at a time, starting with GitHub-only jobs (roadmap, weekly review). Do not port support-triage or Metricool until those MCPs are in ~/.grok/config.toml and proven on a fire. Do not use scheduler_create as the standing trigger.
Legal: treat xAI as Anthropic-class (US, DPA + UK Addendum). Verify the DPA against x.ai/legal/data-processing-addendum the same way clauses 4.3–4.4 were verified (two independent reads). Update sub-processors.md and the published site before product traffic.
Reopen / contradict explicitly: #1189 (routines stay Claude), the parse-pack “do not shrink the model” rule, and clause 3.2 of the sub-processor page.
6. What this research did not do
No live pack-photo bake-off (needs real packs or a fixture set that isn’t in git).
No live ProvenBot citation eval.
No xAI spend quote against current api_usage_log.
No solicitor read of the xAI DPA (same standing as the Anthropic clauses: primary source, deferred legal review).
Did not treat Grok evaluating Grok as independent evidence of OCR quality.
Sources
Repo (13 Sep 2026):supabase/functions/{parse-pack,parse-receipt,parse-recipe,provenbot,admin-api}/index.ts and README.md; WAYS-OF-WORKING.md §3.2a; outputs/docs/ops-automation-design.md; outputs/ops/routine-setup.md and jobs.d/*; outputs/docs/legal-draft/sub-processors.md; outputs/docs/credential-descriptions.yml; outputs/docs/decisions/1092-anthropic-code-review-on-prs.md; outputs/HANDOVER.md (#1189, #1563, #1604, #1593).
28. Card payment providers — who our users have, and how many to accommodate (#433, researched 13 Sep 2026)
Full write-up: outputs/docs/card-payment-providers-research.md (the market evidence, a provider-by-provider verdict table against five needs, the Stripe Connect shape settled far enough to scope, and what could not be verified). Dave's question: which providers will our users already be using to take card payments, and how many should #433 accommodate. Provider capability read from each provider's own developer docs on 13 Sep 2026, not from memory. Headlines:
Who they have. Every 2026 UK guide converges on the same four — SumUp, Zettle by PayPal, Square, Dojo — with a consistent split: SumUp for sole traders and market stalls, Square for shops, Zettle for PayPal households, Dojo for cafés with a fixed till. The only quantified SMB share found (Tuza 2025, via snippet only) puts SumUp at 18.2 %, Dojo 11.3 %, Barclaycard + Worldpay 25.6 % combined; Dojo's own claim is 12.5 % of UK SME card-present acquiring. Guides written for home bakers recommend SumUp and Zettle. A large share take deposits by bank transfer or PayPal with no card provider at all. No survey breaks it down for micro food businesses; the beta-group question #1107 §7 named is still the unlock.
A link is not a reader. A payment link is a card-not-present sale, priced separately by every provider (SumUp 2.5 %, Zettle 1.2 % + 30p, Square 1.4 % + 25p, Stripe 1.5 % + 20p, Stripe Pay by Bank £0). A SumUp merchant paid by a Stripe link still gets the money in their bank; "the provider they already have" matters far less for links than it did for #58's reconciliation, where the data lives only inside that provider. What it drives is friction — a second payout account.
The counter-intuitive finding: SumUp is the likeliest provider our users own and the wrong tool for this feature. Its Checkouts API hosted session expires after 30 minutes and the payments scope is restricted — enabled per application by SumUp via a contact form. Zettle / PayPal POS has links in its app but no API for them, and PayPal's multiparty Invoicing API is "select partners only". Dojo, Revolut, Tide and Monzo hand the merchant a per-merchant secret key, the "paste a full-access key into ProvenBatch" ask #1107 §2.1 already rejected. Worldpay, Barclaycard, Teya, takepayments are contracted acquirers with portal-only pay-by-link.
Only two providers pass all five needs (API link on the merchant's account, OAuth consent, no approval gate, a link that survives until paid, a webhook): Stripe (Connect, Standard accounts, direct charges — and it stores no new class of secret, only an acct_… id, since the platform key plus a Stripe-Account header does the calling) and Square (CreatePaymentLink, OAuth scopes ORDERS_READ/ORDERS_WRITE/PAYMENTS_WRITE, but per-merchant refresh tokens — the new credential class #1107 §5 identified).
Stripe's shape, settled: Standard accounts (Accounts v2 / controller properties for a new platform), OAuth for an existing account or Stripe-hosted onboarding for none, direct charges so the business is merchant of record and carries disputes, "Stripe handles pricing" so ProvenBatch pays no per-account or payout fee, no application fee, a Checkout Session per order with client_reference_id, checkout.session.completed / charge.refunded on a Connect webhook endpoint. The business does nothing in a dashboard — Settings → Connect Stripe → back.
Recommendation: accommodate every provider, integrate one, hold a second behind evidence, never more than two. Layer 0 (effort: low, do first): a per-business "how to pay" block — bank details and/or a link the business made in its own provider's app — printed in the invoice/quote footer band and offered in the order's Ask for payment message; covers SumUp, Zettle, PayPal, Dojo, Revolut, Tide, Monzo and bank transfer at once. Layer 1 (effort: high, #433 proper): Stripe Connect as above. Layer 2 (evidence-gated): Square, only if the beta group both names it as the clear leader and says "make the link for me" rather than "pasting mine is fine". Never a third, never SumUp for links, never anything that needs a pasted secret key.
Unverified, and would change the answer: Tuza's percentages (snippet only; the ranking, not the numbers, is load-bearing); Revolut's API intro (403) and Zettle's API index (404) — verdicts rest on neighbouring pages and neither changes without an OAuth model; Square token lifetimes and link expiry (assumed 30-day tokens with refresh); a tester asking for links in a named provider's account jumps that provider to Layer 2 on the spot.
Sources (13 Sep 2026): docs.stripe.com /connect/payment-links, /connect/accounts, stripe.com/gb/connect/pricing, /gb/payment-method/pay-by-bank · developer.squareup.com /docs/checkout-api, /reference/square/checkout-api/create-payment-link, squareup.com/help/gb · developer.sumup.com /online-payments/checkouts/hosted-checkout, /help · docs.dojo.tech /payments/getting-started, payment-links step-by-step · developer.revolut.com merchant guides · developer.zettle.com, zettle.com/gb/help payment links · developer.paypal.com /docs/multiparty/invoicing · businessofpayments.com (Dojo FY25, Barclaycard) · Tuza Take 2025 (snippet) · mobiletransaction.org pay-by-link roundup (29 Jun 2026) and the 2026 reader guides named in the write-up · foodcore.io home-bakery guide. Full list with what each settled: outputs/docs/card-payment-providers-research.md §6.
Full write-up
From outputs/docs/card-payment-providers-research.md
Card payment providers — who UK small food businesses use, and how many to accommodate (#433)
Researched 13 Sep 2026, against the tree at app v0.369.0 on staging. Provider capability was read from each provider's own developer documentation on that date, not from memory (AGENTS.md #19b); where a page could not be fetched it is marked below rather than filled in. This is the second research pass on #433 (card payment links) — the first, outputs/docs/decisions/433-card-payment-links.md (26 Aug 2026), settled that the real question is "whose account", not "which PSP". Dave's follow-up question on 13 Sep 2026 is the one this answers: which providers will our users already be using to take card payments, and how many of them should we accommodate, given each one is a big job?RESEARCH.md §28 carries the headlines; the provider-by-provider detail is here.
Recommendation, up front
Accommodate every provider, integrate one, and hold a second behind evidence. Never more than two. Concretely, three layers:
Layer
What it is
Providers covered
Size
0 — the business's own link
A "pay this order" field the business fills with a link they made in their own provider's app (SumUp, Zettle, Square, PayPal, Revolut, Tide, Monzo, Dojo…) or their bank details. The app prints it on the invoice / quote footer and puts it in the order's Ask for payment message. No OAuth, no credential, no webhook — the business records the payment by hand, as today
All of them, including bank transfer, which is likely the commonest way a home baker takes a deposit
effort: low
1 — the integrated route
Stripe Connect (Standard accounts, direct charges). The business connects an existing Stripe account by OAuth or opens one through Stripe-hosted onboarding; the app creates the link and hears the payment back on stripe-webhook, recording it on the order like a hand-recorded one
Stripe — the one processor already in the stack (#206). Also the businesses who have no card provider yet
effort: high (the issue's estimate stands)
2 — the evidence-gated second
Square, if and only if the beta group says so. It is the only reader-first provider whose public API lets a third-party app create a durable payment link on the merchant's account through an ordinary OAuth consent
Square
effort: high, plus the new per-tenant-credential class #1107 §5 identified
What is deliberately not on the list, and why (detail in §3): SumUp's API cannot make a pay-later link (its hosted checkout expires after 30 minutes, and the payments scope needs manual activation by SumUp); Zettle / PayPal POS has payment links in its app but no API for them, and PayPal's multiparty Invoicing API is "select partners only"; Dojo, Revolut, Tide and Monzo hand out a per-merchant secret key rather than an OAuth consent, which is the "paste a full-access key into ProvenBatch" ask #1107 already rejected for SumUp; Worldpay, Barclaycard, Teya, takepayments and Lloyds Cardnet are contracted acquirers whose pay-by-link lives in their own portals. Layer 0 serves every one of them.
Why this is the right count. A payment link is a card-not-present sale. It does not need to go through the reader the business already owns — a SumUp merchant paid by a Stripe link still gets the money in their bank account, and the reader is untouched. So "the providers our users already employ" matters far less for links than it did for #58's reconciliation, where the data lives only inside the provider they use. What the existing provider does drive is friction: a second payout account, a second fee line, a second place to look when reconciling. Layer 0 removes that friction for anyone who cares, at almost no cost to us; Layer 1 gives everyone else a link that works; Layer 2 is there for the one provider where the data says the friction is worth a full integration. Every additional deep integration after Stripe adds a stored long-lived third-party credential per business (#1107 §5's finding — ProvenBatch holds none today), a refresh path, a revocation path, a webhook contract and its #279 probes, and a permanent support surface. That cost is linear in providers; the value is not.
1. Who UK small food businesses actually take card payments with
1.1 The market picture
No public survey breaks card-provider usage down for micro food businesses specifically. What exists, in decreasing order of hardness:
Source
What it says
Date
Dojo's own FY25 figures, via The Business of Payments
Dojo claims 12.5 % of the UK SME card-present acquiring market, ~146,000 merchants (flat year on year), £46.2bn volume. Names SumUp, Viva, myPOS, Zettle, FlatPay, Shift4 and Adyen as the pressure on it
Aug 2025
Tuza, Tuza Take 2025 (⚠️ via search snippet only — the page now redirects to a corporate homepage and could not be read at source)
Barclaycard "believed to be losing market share in SME to Dojo", and to Adyen/Stripe/Checkout.com upmarket; Global Payments bought takepayments (2024) and Worldpay (announced Apr 2025); Brookfield bought Barclaycard's merchant business
2024–25
Every 2026 UK comparison guide read (mobiletransaction.org, startups.co.uk, money.co.uk, wise.com, businessexpert.co.uk, promptnews.uk, cardmachineproviders.co.uk, sme-shack.co.uk)
The same four names every time: SumUp, Zettle by PayPal, Square, Dojo. The consistent segmentation: SumUp for sole traders and market traders, Square for shops and growing businesses, Zettle if you already live in PayPal, Dojo for cafés and restaurants with turnover
2026
Guides aimed at our buyer specifically (FoodCore's own How to start a home bakery guide; bakingsubs.com's home-bakery POS guide; hey-dom.com's craft-fair reader guide)
Recommend SumUp and Zettle (£25–£80 readers) for markets, deliveries and collections; Square as the step up
2026
PayPal, asked directly by Payments Dive
Declined to say how many UK merchants use Zettle
—
SumUp press
4 million merchants across 36+ countries; UK is its home market and the segment it markets hardest to is market stalls and sole traders. No UK-only figure published
2025–26
Square
"More than four million sellers" globally; ~9,400 Square Online stores in the UK (storeleads, Q3 2026) — an e-commerce count, indicative only
2026
Teya (ex-SaltPay)
"Over 65,000 businesses" in the UK, contract-based terminals
2026
Reading it honestly: the share figures are for all UK SMBs, where cafés, restaurants, salons and shops with a fixed till dominate; that is where Dojo, Barclaycard, Worldpay and Teya live. Our D-24 buyer — a home baker, a market stall, a caterer, a jam maker — sits in the lowest-turnover band, where the no-contract, no-monthly-fee readers win, and where every guide converges on SumUp first, then Zettle, then Square. Stripe appears in this segment only as the thing behind someone's website or FoodCore's payment links, not as a reader. And a large share of home bakers take deposits by bank transfer or PayPal, with no card provider at all — that is the population Layer 0 (bank details on the document) and Layer 1 (a link for a business with no provider) both serve.
1.2 The provider our users have is a reader; the feature is a link
Every provider above sells its reader on the in-person rate (SumUp 1.69 %, Zettle 1.75 %, Square 1.75 %, Dojo ~1.2 % blended from 2026). A link is a different product with a different rate at the same providers:
Provider
In-person (reader)
Payment link / online, UK card
Link made from
SumUp
1.69 %
2.5 % (0.99 % on the £19/mo plan)
SumUp app / dashboard
Zettle / PayPal POS
1.75 %
1.2 % + 30p by card; 2.9 % + 30p via PayPal
PayPal POS app ("Send link" at checkout)
Square
1.75 %
1.4 % + 25p
Square app / dashboard, or API
Dojo
~1.2 % blended
"remote payments rate" above in-person, custom
Dojo app, or API with the merchant's own key
Stripe
(Terminal, rare here)
1.5 % + 20p; Pay by Bank £0
Dashboard, or API — including for a connected account
Revolut Business
plan-dependent
1 % + 20p domestic
Dashboard / app, or Merchant API with the merchant's key
Tide
—
1.5 % + 9p
Tide app
Monzo Business
—
1.5–3.25 % + 20p (runs on Stripe)
Monzo app
PayPal (no reader)
—
2.9 % + 30p
PayPal.me / invoices
Worldpay / Barclaycard / Teya / takepayments
contract
contract + monthly
their portal, by email
(Fees: mobiletransaction.org's pay-by-link roundup of 29 Jun 2026 and each provider's UK pricing page or 2026 review, read 13 Sep 2026. They move; treat as indicative.)
Two consequences fall out of that table:
A business does not lose anything by taking a link through a different provider from its reader. Card-not-present money is priced separately everywhere, so "my SumUp rate" does not carry over to a SumUp link anyway. A Stripe link at 1.5 % + 20p is inside the same band as every reader-first provider's own link product, and Stripe's Pay by Bank is free.
Every provider our users could plausibly have already gives them a link they can make in thirty seconds on their phone. Layer 0 is therefore not a consolation prize — for a business that already has SumUp or Zettle it is exactly what they would do today with a text message, with the app doing the printing and the bookkeeping prompt.
2. What we would need from a provider, and which ones offer it
For the feature #433 describes — from an order, create a link for a deposit or the balance; hear the payment back and record it on the order — a provider has to offer five things to a third-party application acting for many merchants:
#
Need
Why
A
Create a link on the merchant's account by API
The money must land with the business, not us (decision doc, 26 Aug)
B
An OAuth-style consent, not a pasted secret key
#1107 §2.1 already ruled out asking a customer to paste a full-access key; a per-merchant secret is the same ask under another name
C
No provider approval gate we cannot pass from a cloud session
A "contact us to enable the scope" step is Dave's, is opaque, and may never come
D
A link that survives until the customer gets round to paying
A deposit link is sent Monday and paid Thursday. A 30-minute checkout session is not a link
E
A webhook for the completed payment (and refund)
The order's paid state and Accounts must agree without polling
2.1 The verdict table
Provider
A: API link on merchant's account
B: OAuth consent
C: no approval gate
D: durable link
E: webhook
Verdict
Stripe
✅ Payment Links API / Checkout Sessions with Stripe-Account header (Connect direct charges)
✅ OAuth for Standard accounts, or Stripe-hosted onboarding for a new one
✅ square.link/u/… short URLs; no expiry documented
✅ Payments / Orders webhooks
Hold behind evidence (Layer 2)
SumUp
⚠️ Checkouts API with hosted_checkout.enabled
✅ OAuth (default scopes)
❌ payments scope is restricted — "contact us through our contact form for activation"; a 403 request_not_allowed until then
❌ "A Hosted Checkout session is available for 30 minutes" — then an expired/not-found page
✅ checkout webhooks
Layer 0 only. SumUp's fit is takings import (#1107), not links
Zettle / PayPal POS
❌ API is products, purchases, finance, inventory, image — no checkout or link endpoint (⚠️ developer.zettle.com returned 404/empty on fetch; from the Zettle API index page and the community SDK's endpoint list)
✅ OAuth exists, for those read APIs
—
(app links are durable)
—
Layer 0 only
PayPal Invoicing (multiparty)
✅ create and send invoices for third-party merchants
partner onboarding
❌ "available to select partners only"
✅
✅
Layer 0 only
Dojo
✅ POST /payment-intents → pay.dojo.tech/checkout/{id}
❌ per-merchant secret key (sk_prod_…) from the merchant's developer portal; partner track is "reach out to Dojo support"
⚠️ partner programme by contact
⚠️ 30 days, single use
✅ payment_intent.status_updated
Layer 0 only
Revolut Business
✅ Merchant API orders return a checkout_url (⚠️ intro page returned 403; from the sandbox and hosted-checkout guides and help centre)
❌ per-merchant secret key generated in the Merchant API settings
portal / email pay-by-link; partner APIs are enterprise ISV programmes
❌
❌
✅
—
Layer 0 only
GoCardless / open banking "pay by bank"
(Stripe already offers Pay by Bank on UK Checkout and links at £0)
—
—
—
—
Comes free with Stripe; no separate integration
2.2 The two findings that decide the shape
SumUp is the provider our users are most likely to own, and it is the wrong tool for this feature. That is worth stating plainly because it is the counter-intuitive result. The Checkouts API is built for a webshop's pay-now step: a 30-minute hosted session, a payments scope that SumUp enables per application by hand. Neither is a defect — it is just not a pay-later link. The link a SumUp merchant can make in their own app is durable, which is why Layer 0 exists. SumUp's real integration value for ProvenBatch remains what #1107 found: pulling takings into other_income, on default scopes, once the beta group says enough of them use it.
Stripe Connect with Standard accounts stores no new class of secret. With Standard accounts the platform calls the API with its own key plus a Stripe-Account: acct_… header; the only thing persisted per business is the account id, which is not a credential. Square, by contrast, hands back an OAuth access token and a refresh token per merchant — exactly the per-tenant, long-lived, revocable third-party credential #1107 §5 identified as ProvenBatch's first such secret. That is the single biggest reason Stripe is Layer 1 and Square is Layer 2, over and above Stripe already being in the stack.
3. The Stripe shape, settled enough to scope
The decision doc left "whose account, and which Connect flavour" open. The docs read on 13 Sep 2026 settle it far enough to write the issue:
Question
Answer
Source
Account type
Standard (or its Accounts-v2 equivalent — Stripe now steers new platforms to controller properties / Accounts v2; same properties). Stripe's own example of a Standard-account platform is "a SaaS platform, such as an online invoicing and payment service"
docs.stripe.com/connect/accounts
Existing Stripe account?
Connected by OAuth; Stripe requires OAuth for extensions connecting Standard accounts
same
No Stripe account?
Stripe-hosted onboarding creates one; the business gets the full Stripe Dashboard, so refunds, disputes and payouts are theirs, in their own tool
same
Charge type
Direct charges — the connected account is the merchant of record and carries fraud and dispute liability, not ProvenBatch. Destination charges would make us the merchant of record, which a compliance product for other people's food should not be
docs.stripe.com/connect/payment-links, /charges
Fees to ProvenBatch
None under "Stripe handles pricing": no per-account fee, no payout fee. The business pays Stripe's standard 1.5 % + 20p on a UK card, £0 on Pay by Bank
stripe.com/gb/connect/pricing
Application fee
Optional, fixed amount per one-time payment link (application_fee_amount). Recommend none. We are not a payments business, the D-24 buyer is price-sensitive, and a fee line is a reason to use Layer 0 instead
docs.stripe.com/connect/payment-links
Link vs session
Create a Checkout Session (24 h, one order, one amount, client_reference_id = order id) rather than a Payment Link (persistent, product-based, reusable). A deposit is a one-off amount, not a product. If a longer-lived URL proves necessary, a Payment Link with a single ad-hoc price is the fallback
Stripe API reference
Webhook
checkout.session.completed and charge.refunded on a Connect endpoint (events carry account), handled in stripe-webhook or a sibling function in the house shape; #279 probes in the same commit
docs.stripe.com/connect/payment-links
What the business does
Settings → Connect Stripe → Stripe's own screens (sign in, or the onboarding form) → back to ProvenBatch. No dashboard steps to hand them
same
Everything else in #433's scope (the Ask for payment action, one payments model, tier gate at Standard+, summaryForRange untouched) stands unchanged.
4. How many, and in what order
Step
What ships
Gate
Now-ish, small
Layer 0. A per-business "how to pay" block (bank details and/or a link) on business_settings, printed in the invoice/quote footer band the designer already has (#1064's footer band; today it is free text the business types by hand) and offered in the order's Ask for payment message. Hand-recorded payment unchanged. This closes the FoodCore gap for every provider at once
None. It is the cheapest thing on this page and it makes #433 honest for SumUp and Zettle users
When #433 is picked up
Layer 1, Stripe Connect as scoped in §3
Dave's read of the design record (the issue's first acceptance criterion)
Only on evidence
Layer 2, Square
The beta-group question #1107 §7 already names — "which card reader or payment app do you use?" — plus a second: "would you want ProvenBatch to make the payment link in that account, or is pasting your own link enough?" Build Square only if it is (a) the clear leader among respondents and (b) they answer "make it for me". A fragmented answer, or "pasting is fine", means Layer 0 was the whole feature
Never
A third deep integration; SumUp for links; anything needing a pasted secret key
—
Why not "SumUp too, since they are the most common"? Because §2.2: the API cannot make the thing #433 needs, and the approval gate is SumUp's to open. If SumUp ever ships a durable API-created link on default scopes, revisit — that is the only thing that would move it.
Why not "just Stripe, and no Layer 0"? Because for a business that already has a reader, Layer 0 is what they would do anyway, and it costs a text field. Skipping it would leave the commonest case (SumUp / Zettle / bank transfer) with a worse experience than the rarest (Stripe).
5. What could not be verified, and what would change the answer
Tuza's SMB share figures are quoted from a search snippet; the page has since redirected and the sample and metric (merchants or volume) are unknown. They agree in shape with Dojo's own 12.5 % and with every guide's ranking, which is why they are used, but they are not load-bearing — the ranking, not the percentages, is what §4 relies on.
Revolut's Merchant API introduction (403) and Zettle's API index (404) could not be fetched; their rows rest on the neighbouring guide pages, the help centre and the community SDK. Neither verdict would change if a link endpoint turned up, because both hand the merchant a secret key rather than an OAuth consent.
Square's link expiry and OAuth token lifetimes were not on the reference page read; the usual 30-day access token with refresh is assumed and should be confirmed before Layer 2 is sized.
A beta tester asking for links in a named provider's account is worth more than everything above and would jump that provider to Layer 2 on the spot.
Stripe's Accounts v2 / controller properties are the current recommendation for new platforms; the Standard-account facts above are the same facts under the new names, but the design record should be written against the v2 guide, not the deprecated account-types page.
6. Sources
Read 13 Sep 2026 unless stated.
Provider documentation
docs.stripe.com — /connect/payment-links (charge types, Stripe-Account header, application fees, checkout.session.completed with account, branding), /connect/accounts (Standard / Express / Custom table, OAuth requirement for extensions, liability), stripe.com/gb/connect/pricing (no per-account or payout fee when Stripe handles pricing; 1.5 % + 20p), stripe.com/gb/payment-method/pay-by-bank.
developer.squareup.com — /docs/checkout-api (CreatePaymentLink, quick pay vs order, square.link short URLs, application fees, "all regions where Square payments are accepted"), /reference/square/checkout-api/create-payment-link (ORDERS_READ, ORDERS_WRITE, PAYMENTS_WRITE), /docs/oauth-api/walkthrough; squareup.com/help/gb payment-links article and UK fee schedule (1.4 % + 25p online).
developer.sumup.com — /online-payments/checkouts/hosted-checkout (30-minute session, redirect_url, status via API or webhooks), /help (restricted scopes; payments 403 request_not_allowed; activation via contact form), /tools/authorization/authorization; plus #1107's reading of the OAuth scope table.
developer.revolut.com — hosted-checkout payment-link guide and sandbox setup (Merchant API secret key per merchant account); help.revolut.com merchant API testing article. Intro page: 403.
developer.zettle.com/docs/api (404 on fetch; endpoint list corroborated by the community LauLamanApps/iZettleApi SDK and Shopware's Zettle extension docs); zettle.com/gb/help payment-links article and paypal.com/uk/zettle ("Zettle is now PayPal Point of Sale").
developer.paypal.com — /docs/multiparty/invoicing/ ("available to select partners only"), /docs/invoicing/.
Market and pricing
businessofpayments.com — Dojo: Italy and Spain expansion critical as UK growth slows (26 Aug 2025: 12.5 % share, ~146k merchants, £46.2bn); Barclaycard and Paymentsense tags; April 2025 newsletter (Global Payments / Worldpay; takepayments).
Tuza, Tuza Take 2025 — via search snippet only (SumUp 18.2 %, Dojo 11.3 %, Barclaycard + Worldpay 25.6 %); tuza.co.uk now redirects to tuza.ai.
mobiletransaction.org — 10+ best pay-by-link providers in the UK (29 Jun 2026; the fee table in §1.2) and card machine for small business UK; startups.co.uk, money.co.uk, wise.com, businessexpert.co.uk, promptnews.uk, cardmachineproviders.co.uk, sme-shack.co.uk 2026 reader guides (the recurring four and their segmentation).
foodcore.io/blog/how-to-start-a-home-bakery-uk; bakingsubs.com home-bakery POS guide; hey-dom.com craft-fair reader guide (what our buyer is told to buy).
paymentsdive.com (PayPal declining to give Zettle merchant numbers); sumup.com/en-gb/press; storeleads.app Square UK store counts; mobiletransaction.org Teya review (65,000 UK businesses).
ProvenBatch
Issue #433 and outputs/docs/decisions/433-card-payment-links.md; #58, #1107 and outputs/docs/sumup-integration-research.md (§2.1 on pasted keys, §5 on the new credential class, §7's beta-group question); outputs/docs/stripe-integration-206.md; outputs/docs/invoice-quote-designer-design.md (the footer band); outputs/USER-GUIDE.md ("no reminders or online payment links, just something to hand over or email"); the master competitor sheet (FoodCore payment-links note).
28. Metricool as the growth engine for the 1 Oct launch — tiers, ads, and a loop a robot can run (researched 15 Sep 2026)
Dave's brief: maximise Metricool to grow fast from 1 Oct, as autonomously as possible, Grok orchestrating, every tier and unused service considered, ads included where ROI can be forecast. His framing: paid channels fully reopened, plan against ~£150/month all-in for Oct–Dec, weekly batch approval as the autonomy ceiling. Full write-up: outputs/docs/metricool-growth-research.md; four memos in outputs/gtm/drafts/*-2026-09-15.md.
The live brand corrected the repo twice (read via the connector, 15 Sep). The queue is not drafts — 15 posts 15–27 Sep, autoPublish: true, draft: false, nothing after 27 Sep, so launch week is empty. The brand carries three connected ad accounts: Meta act_1804831247189540 (repo names a different id), Google 9992474342, TikTok Ads 7670576612999725072 (not in the repo).
The claude.ai Metricool connector works on Free: brand settings, scheduled posts, analytics, best-time and plain createScheduledPost all returned; createScheduledPostForReview is 403 (the approval system is Advanced). Metricool's own page says MCP "requires Advanced" — do not pay for Advanced to get what we already have.
Free's real walls are 20 posts/month (help centre, 13 Jul 2026 — the repo's "~50" is wrong) and 30 days of analytics history. October's cadence needs ~24–30 posts ⇒ Starter monthly (~€20, ~£17) from ~28 Sep. Advanced (~€54) buys approval workflow, Looker Studio, API — worth £0 to us (Dave's ✅ in chat is the approval) unless Grok's custom-connector test fails on Free.
30-day baseline (15 Aug–14 Sep): TikTok ~3,100 views from 11 videos (the reach channel by an order of magnitude), Instagram ~540 Reel views from 14 Reels, YouTube ~116 views, Facebook ~1,600 page views and 85 followers. TikTok best times: Wed–Fri 10:00 / 12:00 / 18:00 UK; Instagram returns zeros (too few followers to compute).
Measurement is decision 0. No site analytics, trial_start_click only, no origin on the signup, heard_via retires with /beta. Fix: Cloudflare Web Analytics (D-27, £0, cookieless)
store the already-forwarded utm_*/gclid on the new business (no consent needed) ⇒ cost per started trial by channel, and Google offline conversion import without a tag or banner. The consent-banner route (#1822) stays deferred.
Paid, reopened and modelled. Economics unchanged (allowable CAC £30–100); opt-in trial→paid median ~9–14% in 2026 ⇒ max ~£6–20 per started trial. Per £100 mid case: Google exact-match ~4 trials (£27), Meta boost ~2.5 (£40), TikTok Spark week ~3 (£30) — only best cases clear the ceiling; paid is a learning budget. Plan inside £150: Google £8/day + one £25 boost of a proven Reel + a gated Spark week in November (supersedes 14 Sep (f)/(i) if Dave accepts); Meta conversion campaigns off (need a pixel, fail the ceiling below 25% trial→paid). Gates: continue ≤ £20/trial, kill > £40 after £100 or D7 < 20%, scale only on measured trial→paid ≥ 15%.
Quarter forecast, mid case: ~£400–460 spent ⇒ 32–53 started trials ⇒ 5–8 paying by early January at 15%, payback 4–6 months on July's LTV. "Fast" in this window is tens of trials; the levers beyond that are trial→paid (product) and organic compounding (Flows, TikTok cadence).
The loop: one Grok automation on the digest's proven pattern — Mon 07:30 draft card to #ops (week+1 posts, ≤2 trend ideas, one boost candidate, numbers, ads verdict), Tue 07:30 fulfil on Dave's ✅ via createScheduledPost live, Fri score → beta-pipeline.md §D by PR and a monthly analytics export. £0 of spend ever changed by the routine. Open question: Grok holds no Metricool connector — a ten-minute custom-MCP test at grok.com/connectors decides G1 (Grok executes) vs G2 (the existing Claude Ops: growth routine fulfils). scheduler_create is not a routine; Grok Bot (the agent app) is not required.
Flows (Metricool, 1 Sep 2026, free for now, later a Starter/Advanced add-on): comment keyword → fixed DM with the trial link. The strongest 2026 CTA pattern, and a #12 question — Dave's words, sent by Metricool, triggered by the viewer. Put to him, not assumed.
Worth £0 to us: LinkedIn/X/Pinterest/GBP unlocks (#18), AI credits (our copy rules), reports and Looker Studio (the repo is the reporting layer), SmartLinks (D-12), the web tracker (D-27), Advanced Analytics add-on, social listening (still deferred), Grok Bot.
Not done: the Grok connector test (Dave's account), Keyword Planner (dashboard-only), creating an Autolist on Free (a write), pricing Grok Bot from x.ai (sources conflict).
Capability
Last updated 22 Sep 2026
34. Jev — TypeSafe's System One model, second pass: categorise on entry, everywhere (researched 22 Sep 2026)
Dave's brief: an earlier pass found "relatively few use cases"; repeat it against Nate B Jones's 21 Sep 2026 video on Jev (how LLMs should use Jev alongside their own strengths) and typesafe.ai, with categorisation as the focus. Full write-up, graded and sourced: outputs/docs/jev-system-one-research-2026-09.md. Decided the same day (Dave, 22 Sep 2026): second vendor yes, after 1 Oct; through Cloudflare; recipe lines first, in shadow mode. Filed as epic #2228 (foundation #2229; pilots #2230 recipe lines, #2231 receipt lines, #2232 pack picker, #2233 ProvenBot router; grandchildren #2234–#2245). Updated 24 Sep 2026 (Dave voice, #2448): the 1 October legal/publish gate is lifted early; TypeSafe notice is to go live on Deploy site + promote after land. Shadow switch-on remains a separate PR. The earlier pass is not in the tree, the issues, the branches or the artifact gallery, so this one starts from scratch.
What it is. A classifier with calibrated probabilities: state in, typed answers out — Choice (one of up to 255), Score (2–10 described levels), Noul (a probability a statement is true) — all questions in one call evaluated in parallel, 70–500 ms, $0.042 per million input tokens, output free, 64k context, text only, jev-1.13.0. Cannot write, count, compare dates or see an image; can return the wrong valid option; reports confidence and is trained so it means something. Also reachable through Cloudflare as typesafe/jev (a pass-through, 32k).
The division of labour the video and the vendor both state: code keeps control; Jev makes the narrow typed judgements; the LLM reads photos, writes and reasons. The video's test is "complicated inputs and simple outputs", and its instruction is to look for where a decision should sit, not where an LLM already is. (The captions were unreachable from the container; the episode's published show notes were used and the doc says so.)
The inventory that changes the answer (doc §3): ProvenBatch makes ~40 bounded decisions over messy input and only 4 LLM calls, none Jev-shaped (vision or generation) — which is why a pass starting from the LLM calls found little. 14 are exactly Jev-shaped, today all similarity scores with once-measured thresholds (bestCatalogMatch ≥ 0.5, SUGGEST_SIM_FLOOR 0.29, packNameEvidence), a 60-word keyword router (isComplianceTurn), or a default the user must notice ("Supplier products" on every receipt line).
Ranked fits (doc §4): A receipt lines — category + catalogue match + is-it-a-product in one call per receipt; B the pack "one you already have?" picker and duplicate sweep (pack_parse_events already holds the ground truth); C recipe lines → ingredients; E the ProvenBot model router; G the allergen scan as a second opinion into the Accuracy check, never the decision; then file triage, feedback triage at submit, FSA recall borderlines.
The categorisation rule (doc §5): categorise on entry, everywhere; pre-fill at ≥ 0.9, chip at 0.5–0.9, leave alone below, a different bar per field by the cost of being wrong; never ask it to name — offer candidates from the matchers we have; policy lives in the option descriptions as a curatable vocabulary; sweep the catalogue nightly; log the answer beside what the user chose from day one — that is the eval.
Money (doc §6): a receipt's Jev decisions ≈ £0.0002, about 2% of the Haiku read; the same on Haiku would double the receipt's cost. No saving on today's four calls; the one direct saving is the router, ~5–10 US cents per misrouted turn, measurable with anthropic-cost. Side finding: provenbot and admin-api price constants are older generations' rates (display drift only; the cost job reconciles).
The gate is not technical (doc §8): a second product-runtime AI vendor after decision 1939 took it to one; a new US sub-processor — DPA with EU SCCs + UK Addendum, no training on inputs, retention "as long as necessary", ZDR enterprise-only — with our clause-3.2 in-app notice before any customer row is sent, shadow included; a credential-inventory entry and a slug kept out of AI_SLUGS; one system-one edge function with question sets versioned in the repo; today's heuristic as the fallback behind every call.
Pilot order as decided: recipe lines (#2230) in shadow mode for two weeks, thresholds from the accuracy-per-band table, then pre-fill on; receipt lines, the pack picker and the router follow. Nothing before 1 Oct 2026; the foundation child (#2229) lands first.
Full write-up
From outputs/docs/jev-system-one-research-2026-09.md
Jev (TypeSafe's System One model) — where it fits ProvenBatch, second pass (researched 22 Sep 2026)
Correction, 29 Sep 2026 (#2666, #2718): the 1 Oct 2026 public trial mentioned below is postponed with no new date; sign-up stays invite-only (decision #2666). Read "the 1 Oct public trial" as "the public opening". Findings are as researched on 22 Sep.
Researched 22 Sep 2026. Dave's brief: an earlier pass "indicated relatively few use cases … I still think there is vast potential to use this model with our app. Repeat the research, but this time reference [Nate B Jones's video] as to the sorts of use cases Jev has — in particular how the host says LLMs should look to utilise Jev alongside their own strengths — and the Jev site, typesafe.ai. Categorisation is key here. We must be categorising so much in the app. How can we have Jev help us do this fast and cheap to really accelerate our users' experience and maybe even save us on costs at the same time?" Sources: the video's published show notes (the captions were unreachable from this container — §2 says exactly what was and was not read), every page of docs.typesafe.ai that bears on the question including nine cookbooks, TypeSafe's privacy policy and DPA, Cloudflare's model page, eight independent write-ups and two independent tests, and a file-and-line inventory of every place ProvenBatch categorises, matches, routes or ranks, built against main the same day. The earlier pass could not be found — not in the tree, the issue tracker, the branches or the artifact gallery — so this one starts from scratch and, in §7, says why "few use cases" was probably the right answer to a narrower question.
How to read the grades.[verified] = the primary page, spec or document was fetched and read on 22 Sep 2026; [reported] = a secondary write-up or a search extract, not the primary source; [inference] = our reasoning from the evidence; [repo] = a fact from this repository or its recorded production measurements, file and line given.
Decided the same day (Dave, 22 Sep 2026): a second product-runtime AI vendor is acceptable, after 1 October 2026; access goes through Cloudflare (typesafe/jev on Workers AI); recipe lines are the first pilot, in shadow mode. Filed as epic #2228 with a foundation child (#2229) and one child per pilot (#2230 recipe lines, #2231 receipt lines, #2232 pack picker and duplicates, #2233 ProvenBot router), each with a shadow and a switch-on grandchild (#2234–#2245). No vendor account is opened and no customer row is sent before the foundation child lands. §9 records the calls as made.
Updated 24 Sep 2026 (Dave voice, #2448): the 1 October legal/publish gate is lifted; Deploy site no longer waits for 1 Oct for the TypeSafe notice. No customer row until that notice is live. Shadow flags stay off until their own switch-on issues.
0. Headlines
Jev is a classifier, not a small LLM, and that is the whole point. [verified] State in, typed answers out — one of N (Choice), a place on a rubric (Score), a probability that a statement is true (Noul) — every question evaluated in parallel, 70–500 ms, $0.042 per million input tokens with output free. It cannot write a word, cannot count, cannot compare dates, cannot see an image, and can return the wrong valid option. It also reports how sure it is, and is trained so that number means something.
The division of labour the video and the vendor both state: code keeps control of flow, arithmetic and side effects; Jev makes the narrow, typed judgements over messy input; the LLM does what only it can — read a photo, write a sentence, reason across hops. The video's test for a Jev-shaped problem is "complicated inputs and simple outputs", and its instruction is to ask where intelligence should appear throughout a workflow, not where an LLM is already called — because the price of a classification has fallen far enough to revisit every decision that was left to the user for cost reasons [verified — the episode's show notes; docs.typesafe.aiHow to build].
Asked that way, ProvenBatch has around forty bounded decisions over messy input and only four LLM calls [repo — §3]. None of the four is Jev-shaped (two are vision, one is generation, one is both), which is why a pass that started from the LLM calls would find little. Fourteen of the forty are exactly Jev-shaped, and today every one of them is a hand-written similarity score with a once-measured threshold, a keyword list, or a default the user has to notice and change.
The categorisation answer (§5): categorise on entry, everywhere, one call per record asking every question at once, the answer pre-filled at ≥ 0.9 confidence, offered as a chip between 0.5 and 0.9, left alone below — with a different bar per field according to the cost of being wrong. Never ask it to name a thing: our existing matchers become the candidate generators and Jev the judge. Put the policy in the option descriptions, where an admin can curate it like allergen_terms. Sweep the whole catalogue nightly for pennies.
Ranked fits (§4): A receipt lines — category, catalogue match and is-it-a-product in one call, the highest click count and ground truth for free; B the "one you already have?" pack picker and the duplicate sweep, where pack_parse_events already records what the user did next; C recipe lines to ingredients; E ProvenBot's model router, today a 60-word keyword list that sends "my label printer won't connect" to Opus at effort high; G the allergen scan as a second opinion into the existing Accuracy check, never the decision; plus file triage, feedback triage, recall matching and a handful of small ones.
Money (§6): a receipt's worth of Jev decisions is about £0.0002 — 2% of the Haiku read it follows; the same decisions on Haiku would double the receipt's cost and on Sonnet treble it. Jev does not save money on today's four calls; it makes the next fourteen decisions affordable at all. The one direct saving is the ProvenBot router, of the order of 5–10 US cents per misrouted turn [inference], measurable before and after with the anthropic-cost job that already runs daily.
What does not move (§7): pack, receipt and recipe reading (Opus and Haiku stay), any writing, the label engine and the allergen decision, anything with arithmetic or dates, HACCP judgements, the data-quality rules.
The costs of yes (§8) are the real gate, not the technology: a second product-runtime AI vendor three weeks after decision 1939 took it to one; a new US sub-processor with our own clause-3.2 in-app notice due before the first customer row is sent — shadow mode included; a credential and a slug kept out of the Anthropic ceiling; one edge function with the question sets versioned in the repo; a shadow log that is the eval; and today's heuristic as the fallback behind every call. Jev is reachable through Cloudflare as typesafe/jev, which removes the new-account half of that list but not the sub-processor half.
1. What Jev is, in the terms that matter for a decision here
A classifier with calibrated probabilities, not a small language model. [verified — docs.typesafe.ai System One, Primitives, Confidence, Models, API reference; typesafe.ai/blog/introducing-system-one-models-and-jev] TypeSafe AI, Inc. (US; founder Diogo Almeida, an InstructGPT co-author; launched 15 Sep 2026 with $40M) calls Jev a "System One model": you send state (a string, a JSON object or an array of text, up to 64k tokens) plus one or more typed questions, and it returns typed answers in one parallel pass, 70–500 ms end to end. It cannot write a sentence. There are exactly three question types:
Primitive
Asks
Returns
Limits
Choice
pick one option from a named set, each option described
choice, a probability per option, confidence 0–1
up to 255 options per question; chain questions for deeper taxonomies
Score
where on an ordered rubric this sits
score (probability-weighted mean across the levels, so 1.43 is a legal answer), a probability per level, confidence
2–10 levels, each a described situation, not a degree
Noul
is this statement true of the state
one probability 0–1 that the answer is yes
no separate confidence; threshold it in code
Every question in a request is evaluated independently and in parallel against the same state, so asking twelve questions costs tokens but almost no extra time (the vendor's own 13-question benchmark: one batched call was 12.2× cheaper and 10.0× faster than thirteen sequential ones) [verified — cookbook Parallel questions].
Price and speed [verified — docs.typesafe.ai/models, the launch post]: $0.042 per million input tokens; output is free. Rate limits 250k tokens/s and 1,200 requests/min. Context 64k per request, 32k for state plus the longest single question. English is the training language; other languages work with lower accuracy. Text only — no images, audio or video. The model is jev-1.13.0; jev-latest is the SDK default; pin the version if you tune thresholds against it. Python (typesafe-sdk) and JS/TS (@typesafe-ai/sdk) SDKs, or one POST https://api.typesafe.ai/v1/systemone with a bearer key — the HTTP shape is small enough that a Deno edge function needs no SDK at all.
Confidence is the second axis, and it is the reason to use this rather than a prompted LLM. [verified — docs.typesafe.ai/confidence] Confidence collapses the probability distribution to one number: a single peak is high, a flat spread is low. The vendor's guidance is three bands — act automatically above about 0.9, confirm with the user in the middle, route below about 0.5 to a person or a reasoning model — and, crucially, different actions in the same system get different thresholds according to the cost of being wrong. TypeSafe trains for this ("Reinforcement Learning for Calibrated Decisions"), which is why they argue a Jev probability means something a prompted LLM's "how sure are you" does not. The independent readings are more careful: no calibration curves or Brier scores have been published, the 67.8% headline accuracy is agreement with two frontier models rather than human ground truth, and on the hardest of the vendor's four workflows (invoicing) Jev scored 61.8% against 74.7% for the comparable frontier model [reported — pearpages.com 16 Sep 2026; firecrawl.dev]. Two outside tests are worth recording: Every's editorial checks found Jev caught six of seven planted defects where Claude Fable 5.1 caught all seven, at a median 0.35 s a passage against 8.83 s and roughly 580× lower cost [reported — firecrawl.dev citing every.to]; a 60-agent simulation ran 13,200 decisions for $0.35 against a simulated $37.64 on a frontier model, and its one real failure was a vaguely-worded criterion, fixed by stating the condition ("only when hunger is above 50") in the option text [reported — tylerfolkman.substack.com].
The vendor's own list of what it is bad at [verified — docs.typesafe.ai/model-jaggedness/jev-1.13] is the list that shapes every design in §4 and §5:
Literal reading — it answers the question as written, not as meant. Put the exact condition and the boundary cases in the option descriptions.
Numbers — it does not count or compute. Any arithmetic (pack size × unit price, per-100g maths, VAT) stays in code.
Dates — it cannot order dates or compute intervals. Extract the parts as Choices, compare in code.
Indirection — double negatives and multi-hop questions lose accuracy. Name the field.
Large irrelevant state — accuracy falls as the state fills with things the question does not need. Filter in code first ("retrieve, then judge").
Adversarial content — an instruction hidden in the state can move an answer. A receipt or a scraped page is untrusted input.
Contradictory instructions and criteria — align them.
No structural invariants across questions — P(yes) on one Noul and P(no) on its inverse need not sum to one; do not transfer a threshold from one primitive to another.
No generation — it will not name a thing it has not been offered. Extraction is done by offering candidates (a regex's spans, a catalogue's rows) and letting it choose.
Two other facts that bound the design. "Cannot hallucinate" means cannot return a value outside the schema; it can and does return the wrong valid option, so a Jev answer is never a source of truth for a label — it is a suggestion the user confirms or a gate that decides which slower path runs [verified — the vendor says so in the jaggedness page and in its own caveats on the launch post]. And Jev is reachable through Cloudflare as typesafe/jev on the Workers AI binding (env.AI.run('typesafe/jev', …)) and the account REST endpoint, as a pass-through to TypeSafe's infrastructure, not Cloudflare-hosted inference, with a 32k context on that route [verified — developers.cloudflare.com/ai/models/typesafe/jev/]. That removes the new-account and new-secret half of adopting it (billing lands on the Cloudflare account we already have) but not the sub-processor half: the state still reaches TypeSafe (§8).
2. How the video, and the vendor, say an LLM should use Jev
The video. Nate B Jones, AI News & Strategy Daily, 21 Sep 2026, 33 minutes: on YouTube as "Why Developers Are Losing Their Minds Over AI That Can't Write" (youtu.be/tYugqJ9YytQ; the live title was "What is Jev? The AI that Can't Talk Back (and why that's a good thing)"), published the same day as the podcast episode "You cannot tell which parts of your software should stop calling an LLM. My Jev guide has a prompt that scans your projects and names them" [verified — YouTube oEmbed for the title and channel; Acast episode page for the date, length and show notes]. ⚠️ The captions could not be pulled from this container — YouTube served its bot wall to the watch page, the player API, the transcript API and every transcript mirror tried, and the companion Substack guide is paywalled below its headline. What follows uses the episode's published show notes and the vendor's own "how to build" page, which state the same division of labour; where the video is quoted, the quote is from the show notes, not from the audio. If Dave wants the exact wording, the guide's "prompt that scans your projects and names them" is the one thing in it this research could not reproduce, and §3 of this document is our hand-made equivalent of what that prompt would return.
The episode's thesis, from its show notes: "complicated inputs and simple outputs define a useful class of problems"; the question to ask of a codebase is where intelligence should appear throughout workflows, not where an LLM is already called; the three parts of a modern system are classifiers, generative models, and standard code, and they interconnect rather than compete; and the reason to look now is that the price of a classification has fallen far enough that "decision-making can shift from humans to automated systems" and builders should "revisit architectural decisions made expensive by prior constraints and experiment with deploying classification throughout complete workflows." The worked examples named are support routing, tax-document analysis, research prioritisation, agent orchestration and dynamic spreadsheets [verified — Acast show notes].
The vendor's version of the same rule [verified — docs.typesafe.ai/concepts/how-to-build-with-system-one, /concepts/use-case-map, /patterns/*]: "build a normal software workflow and insert System One only where AI is needed." Code keeps control — control flow, deterministic rules, arithmetic and side effects stay in code; the model is asked only for common-sense judgements over unstructured data. Decompose the state — send the fields the question needs, not the record. Decompose the question — "is this spam?" becomes three Nouls (asks for credentials; sender mismatches domain; announces an unexpected reward) combined with weights in code, so each judgement is inspectable and the weights are yours to tune. And the LLM is used for what only it can do: read an image, write a sentence, explain a decision, reason across many hops.
The four documented patterns, and the cookbooks that turn out to be direct templates for our problems:
Pattern
Shape
The ProvenBatch problem it fits
Speculative fan-out
ask every question you might need in one call; ignore the answers the first answer makes irrelevant
one call per receipt line answering category, packaging-or-ingredient, VAT-likely, is-this-a-fee — with the later answers used only if the first says "ingredient"
Confidence-gated routing
the answer says what; confidence says whether to act, confirm, or escalate
pre-fill a dropdown at ≥0.9, show it as a suggestion in the middle, leave it blank and ask below 0.5 — a different bar for an allergen than for an expense category
Intent routing
classify the request, then hand it to pure code, a specialist LLM prompt, or a person
ProvenBot's first hop; feedback and support triage
Composite scoring
score several independent dimensions, combine with weights in code
"is this pack photo good enough to read" from three Nouls; feedback urgency
Hierarchical classification (cookbook)
beam search over Choice probabilities down a taxonomy of 12,500 retail product types
an ingredient's type and family; an expense's accounting category
Entity alignment (cookbook)
one three-level Score — different / closely related, hand to a curator / the same — over 450 candidate pairs from two beer catalogues: 80% settled as different, 8.9% merged, 11.1% to a person
regex finds the candidate spans, Jev picks the one the question means, code copies it verbatim — it "cannot invent a value or transpose a digit"
which of the numbers on a receipt line is the quantity, the unit price and the line total; which date is the invoice date
Structured-data-extraction cascade (cookbook)
cheap extractor → Jev verifies each field with seven Nouls → escalate only the fields that fail to a reasoning model
a per-field check on what parse-receipt and parse-pack return, before the user sees it
Classification using confidence (cookbook)
75 industry groups; at ≥0.9 report the group, below it report the parent division — useful answers rose from 65% to 80% with one request per document
assign a fine category when sure, the coarse one when not, never nothing
Guardrails for LLMs (cookbook)
screen what goes into and out of an LLM with hazard Nouls and a severity Score
the ProvenBot compliance answer, and every scraped page fetch-page brings back
The use-case map's own taxonomy of "task shapes" — classification, detection, scoring, routing, search, retrieval, ranking, verification, feature extraction, structured extraction — is the checklist §3 was built against.
3. What ProvenBatch decides today, and how — the inventory the video asks for
Built by reading outputs/bakery-app-site/src/part-1.js and part-2.js, every supabase/functions/*/index.ts, outputs/scripts/recall-match.js, the ops registry and the migrations' CHECK constraints on 22 Sep 2026 [repo]. The question asked of each place was the video's: is this a complicated input with a simple output? — not does this already call a model? The two questions give very different answers. Only four places call a model today, and none of them is Jev-shaped (two read images, one writes prose, one does both). Around forty places make a bounded decision over messy input, and fourteen of them are exactly what Jev is for.
#
Decision
Where
Input
Output shape
Decided today by
Jev-shaped?
1
Which expense category a receipt line belongs to
expenseLineCategorypart-2.js:1815
line description, supplier
1 of 5: Supplier products / Packaging / Equipment / Fees / Other (EXPENSE_CATEGORIES:1664)
precedence: catalogue match → supplier default → "Supplier products" always
Yes
2
Which catalogue item a receipt line is
bestCatalogMatch:6729, catalogMatchScored:6696
description vs supplier-scoped item names
pick one of N, or none
Jaccard token overlap ≥ 0.5, substring bumps to 0.6
Yes — retrieval exists, judgement is crude
3
"Is this the same thing you buy elsewhere?"
crossSupplierSuggestion:6791
description vs other suppliers' items
one chip, or nothing; never auto-applies
evidence gate + similarity floor 0.29 (measured over 20 pairs)
Yes
4
Is this receipt line a product, or a fee / deposit / discount / tender line
regex over a curated vocabulary; additive, verified:false, human confirms
Yes as a second opinion, never as the decision (§4-G)
18
Which mandatory statement a declaration triggers
statementTriggerHits :13779
declaration text
multi-label over mandatory_statements
regex over statement_trigger_terms
Same posture as 17
19
Is the pack's ingredients panel legible; is this text a declaration at all
parse-packlegible, clean() :93
the photo
boolean
Opus, inside the vision read
Partly — Jev cannot see the photo, but can judge the text it returned
20
Which fields of a read are wrong
nothing — the user finds out at confirm
the returned JSON
per-field flag
nobody
Yes — the SDE-cascade cookbook
21
Is a feedback item a bug, a question, a request; how urgent; about an allergen
submit-feedback :245 (one fixed label); support-triage routine next morning
the message
type, urgency, flags
nothing at submit; an LLM routine daily
Yes
22
Does an FSA recall match a supplier product
recall-match.js scorePair :370
alert product vs catalogue
definite / borderline / none
an eight-tier ladder of thresholds; Haiku writes one explaining sentence for borderline
Yes for the borderline judgement; the sentence stays Haiku
23
Which data-quality class a finding is
dataQualityFindings :19137
catalogue state
blocking / silent / hygiene
deterministic rules
No — nothing unstructured in it
24
Which label template, date type, format
labelTemplateFor :12141, labelSafety
nation, format, sourcing state
one template; blocking / warn
deterministic; DB rows win
No, and must stay no (#15, #16)
25
Which catalogue rows answer a ProvenBot search
pbSearchCatalogue :19728
query vs names
ranked 25
token tiers exact / prefix / typo
Partly — a rerank of the 25 for a natural-language query
26
HACCP hazard type, critical limits
haccp.js:17, :315
the step
biological / chemical / physical / allergen
the user, by design; the app never suggests a limit
No, by design (§7)
27
Business type at signup
BUSINESS_TYPES :693
free text
9 types + other
the user; ProvenBot can propose
Yes, trivial
28
Which task to suggest today
suggestedTasks :7726
live state
list
deterministic
No
29
Has a customer journey got slower since a release
journey-watch/rule.js
timings
yes / no
statistics with a bootstrap CI
No
30
Which issues can run in parallel; which effort label
create-issues skill, modelroute2112
issue text
effort low / medium / high; isolation
the filing session
Yes, but dev-side, not product
Rows 1–3, 7–10 and 12 are the ones Dave means by "we must be categorising so much in the app" — and today every one of them is either a hand-written similarity score with a threshold someone measured once, or a default the user has to notice and change. Row 13 is the one that is already spending Opus money on a keyword list.
Volumes [repo — caps and measurements, not traffic]: receipts are capped at 20 reads a day per user server-side and 50 client-side; a pack is read once for life (120/day cap; 16 real reads in the 24 Aug measurement, £0.0161 each); ProvenBot is 120 calls a day with a published 30 / 150 / 500 turns a month by tier; recall matching runs nightly over every tenant's catalogue. None of this is high-volume in Jev's terms — the point is not that we have millions of decisions, it is that at this price the per-decision cost stops being a reason to leave a decision to the user.
4. Where Jev fits — ranked, with the shape of each design
Each entry gives the question set, the gate, what the user sees, and what does not change. The ranking weighs how many user clicks it removes, how often it fires, and how safely a wrong valid answer fails. Costs are from §6.
A. Receipt lines — category, catalogue match, and "is it even a product" in one call [highest value, cleanest fit]
Today [repo]: parse-receipt (Haiku, vision) returns the lines; the client then runs a Jaccard overlap against the supplier's own catalogue (≥ 0.5 wins), falls back to the supplier's default category, and otherwise writes "Supplier products" for every line. A new business with an empty catalogue therefore gets the same category on every line of every receipt, and #2162 (landed 21 Sep) is about making the correction select visible enough that they notice. The cross-supplier chip (SUGGEST_SIM_FLOOR = 0.29) never auto-applies because the similarity score is not a probability — the comment above it says as much.
With Jev [inference, on the vendor's fan-out and entity-alignment patterns]: after the read, one request per receipt, state = {merchant, supplier_default_category, lines: [{i, description, qty, total}], candidates: [top 6 catalogue names per line from the existing matcher plus the supplier's most-bought items], categories: the five, with a one-line "what / not_for / examples" each}. Per line, three questions asked together:
category_i — Choice over the five, descriptions written so that "Packaging" says boxes, bags, jars, labels, tape bought to sell product in, "Fees" says delivery, card, subscription, licence, bank charge, "Equipment" says a thing used more than once, "Other" is the honest escape. The literal-reading rule (§1) means the descriptions carry the policy, not the prompt.
item_i — Score over three levels, the beer-catalogue shape: a different product / closely related, a person should look / the same product — over the offered candidates, plus a Choice naming which candidate if level 2. It cannot invent a catalogue item; it can only pick one we offered.
is_product_i — Noul: this line is a purchased item, not a total, VAT, deposit, discount, loyalty or tender line. A text second opinion on what Haiku decided from the image.
Gate [vendor guidance]: category pre-filled at confidence ≥ 0.9, shown as a suggestion chip at 0.5–0.9, left on today's default below 0.5. Catalogue match: level 2 links the line exactly as #69 does today; level 1 renders the existing suggestion chip; level 0 nothing. is_product below 0.3 collapses the line into the "not a product" state the review screen already has.
What the user sees: the review screen arrives with categories already right and matches already made; the select from #2162 is still there, and now it is for the exceptions. What does not change: nothing is saved without the user confirming the draft, exactly as now; the VAT read stays Haiku's; the arithmetic stays code.
Bonus questions in the same call, free: vat_plausible_i — a Choice over the tenant's VAT_RATES from the description alone (most food is zero-rated; hot food, confectionery, drinks and non-food are not) compared with the printed code Haiku read, flagging disagreement; and packaging_or_ingredient_i for the Packaging catalogue. Volume: bounded by the receipt caps. Shadow-mode ground truth exists already: the category the user ends up saving.
B. "One you already have?" after a pack photo, and the duplicate sweep [high value; telemetry already measures it]
Today [repo]: packScanMatches ranks by a barcode hit or evidence × 2 + brand, returns three, and pack_parse_events records match_offered ∈ {barcode, name, none, unmatchable} and match_action ∈ {created_new, used_top, used_other, used_list} with the margin. The nightly duplicate grouping (#1984/#1985) uses the same evidence gate and a 0.29 floor, then union-find.
With Jev: keep the barcode path (it is exact) and the evidence pre-filter (it is the retrieval half); replace the judgement over the surviving candidates with the three-level Score from the entity-alignment cookbook, plus three Nouls the cookbook uses as supporting detail — same brand, same product name, same pack size — so the verdict sentence in packSuggestionHtml can say why. For duplicates, the same Score per candidate pair, with level 1 becoming the "possible duplicate" data-quality finding and only level 2 proposing a merge. Gate: level 2 above 0.85 offers "use this one" first; level 1 offers it in the list; nothing is linked without the tap, as now. Measure before switching: pack_parse_events already holds what the user did after each offer, so a week of shadow answers gives accuracy per confidence band with no new instrumentation.
C. Recipe lines → ingredient, and pasted text → lines [high value at onboarding]
Today [repo]: bigram Dice ≥ 0.5 picks the ingredient, runnerUpIngredientMatch (#312) names one alternate, and a pasted line with no digit and ≤ 40 characters becomes a section heading. parse-recipe (Haiku) does the same on photos and PDFs.
With Jev: per pasted line, a Choice over the tenant's ingredient names (chained hierarchically above 255 — by family first) with an explicit none of these, and a Choice for the line's role: ingredient / method step / section heading / equipment / serving note / oven temperature. On the AI path, the same call verifies Haiku's split. Gate: pre-select at ≥ 0.9, offer the runner-up from the probability list (a real second-best, not a heuristic one) below. The unit maths stays in parseQuantity and unitMap.
D. What the user just dropped [medium; onboarding only]
classifyGetStartedFile reads a filename regex and a CSV header row. For a spreadsheet, CSV or PDF, send the header row and first three rows (or the first page's text) and ask one Choice over the seven kinds plus, for a spreadsheet, a Choice per column over the target fields — which is the column-mapping step onboarding does not yet have. For a photo Jev has only the filename; the heuristic stays. The dropdown the user can override (getStartedKindOptions) stays.
E. ProvenBot's first hop — intent, complexity, and the model router [the one with a money number]
Today [repo]: isComplianceTurn() is a 60-word list. "My label printer won't connect" contains label and goes to Opus at effort high; "can I sell brownies at a market without registering" contains none of the words and goes to Sonnet at effort low. The comment says the keyword test was chosen deliberately over a model call — the reason was cost and latency of the model call, which is the exact thing that changed.
With Jev: one request on the message (and the last two turns), three questions: intent — Choice over compliance question / how-do-I-use-the-app / look something up in my catalogue / costing or pricing maths / small talk / other; compliance — Noul with the current keyword list folded into its criteria so the floor never gets worse; complexity — Score over three situations (a lookup / needs judgement / unusual, may need a person). Route: compliance or low confidence → Opus high (as now); lookup → the search_catalogue tool path on Sonnet; how-to → Sonnet with search_user_guide pre-called. Gate: below 0.5 confidence, route to Opus — the expensive path is the safe default, so the router can only save money, never quality. Latency adds ~150 ms before the first token. The open_screen.tab free-text field becomes a Choice over the screen list in the same call.
F. Verify what the readers return, per field [medium; the vendor's SDE cascade, without the escalation]
Jev cannot see the receipt or the pack; it can judge the text Haiku and Opus returned. Per field, Nouls of the kind the cookbook lists: the merchant field is a trading name, not an address or a slogan; this "item" is a product line; this declaration text reads as an ingredients list, not a nutrition table or cooking instructions; the emphasised words all appear in the declaration. A failed field is flagged on the review screen, not corrected. What it cannot catch, stated plainly: a fluent, wrong declaration — the labelling incident RESEARCH §26 names — because that text is internally consistent. Opus stays on packs; the "do not shrink the model" rule in parse-pack stands.
G. The allergen scan — a second opinion into the existing Accuracy check, never the decision [compliance-sensitive; design first]
Posture, fixed before any design [repo: #15, #16, bakery-manager-labelling-solution-design.md, DESIGN-LANGUAGE.md §Bulk select & edit]: labels are derived from confirmed sourcing rows; the regex plus the curated allergen_terms vocabulary plus a human tap is the chain of truth; autoScanSourcing only ever adds with verified:false. Nothing here changes that.
With Jev, as a checker: fourteen Nouls over the declaration text — this declaration lists an ingredient derived from milk, … — each with the FIC Annex II derivative list in its criteria. Compare with scanAllergens. Where Jev says yes and the regex said nothing, raise the existing allergen_unverified-style finding with the evidence phrase; where the regex says yes and Jev says no, do nothing (the regex is the floor). Run it as a nightly map-reduce over every declaration (§6 — pennies) rather than in the save path, and feed the disagreements to the admin who curates allergen_terms: a word that trips Jev and not the regex is a candidate term. The "wrong valid answer" failure is contained by construction — it can only add a question for a human, never remove a row or change a label.
H. Feedback and support at the moment of submission [small volume, high leverage at launch]
submit-feedback applies one label; a Grok routine triages next morning. One request at submit: typeChoice (bug / question / feature request / praise / billing / data or privacy / other), urgencyScore (three situations), compliance_flagNoul (mentions a wrong allergen, label or declaration), frustrationScore. Write them as GitHub labels on the issue the function already creates. The routine and Dave then read a sorted queue; a compliance flag can page. No reply is ever generated — #12 is untouched.
I. FSA recall matching — the borderline judgement [nightly batch; free]
recall-match.js is an eight-tier ladder of thresholds ending in definite / borderline / none, and Haiku writes one sentence for the borderline rows without deciding. Keep the ladder as the candidate generator; replace the borderline judgement with the same three-level Score as B over {alert product, brand, pack; supplier product, brand, pack}; keep Haiku for the sentence (Jev cannot write it). The RECALL_MATCH_KINDS enum does not change.
J. Smaller, still worth listing
Supplier from merchant name (row 6): a Choice over the tenant's suppliers plus new, with chain_profiles aliases in the descriptions.
Product category and ingredient type (categories is user-created; type ∈ ingredient / resale / both): a Choice over the tenant's own list plus new at create time.
Business type at signup from the business name and the free-text field: a Choice over the nine.
ProvenBot catalogue search: rerank the top 25 for a natural-language query with one Noul per row (this row answers the query) — the rerank cookbook took top-1 from 5% to 18% on legal passages; ours is a smaller, cleaner problem.
Effort labels and isolation notes when filing issues: a Score on the issue body; dev-side, so it does not touch the vendor question.
5. Categorisation, answered directly
Dave's question was how can we have Jev help us categorise fast and cheap, to accelerate the user's experience and maybe save on costs at the same time? The answer the video, the vendor's docs and the inventory agree on:
Categorise on entry, everywhere, and stop asking. Every one of rows 1–3, 7–10 and 12 is a field the user currently fills or corrects. At ~150 ms and a hundredth of a penny, the category can be filled before the user reaches the field — a debounced call as the description is typed, or one call as the read returns. The select stays; it becomes the exception path.
Confidence decides what the user sees, and the bar is set by the cost of being wrong. Three states, everywhere, so the app is consistent: filled (≥ 0.9, act), suggested (0.5–0.9, a chip to tap), empty (< 0.5, today's behaviour). An expense category wrong is a tax box wrong at year end — filled at 0.9 is fine. An allergen wrong is a labelling incident — never filled, only ever a question for a human (G).
Never ask it to name a thing. Offer candidates and let it choose. Every matcher we already have — Jaccard, evidence gate, barcode, alias tables — becomes the retrieval half; Jev is the judgement half. That also keeps every Choice under 255 options and keeps the state small, which is the vendor's own first rule for accuracy.
Many questions, one call. Per record, ask everything at once (category, match, is-it-a- product, VAT plausibility, packaging-or-ingredient) and let code discard what the first answer makes irrelevant. The cost is tokens; the latency is one round trip.
Put the policy in the option descriptions, not in a prompt. Jev reads literally. "Fees: delivery, card, subscription, licence, bank charge" is a policy the admin can edit and a test can pin, in the way allergen_terms and chain_profiles already work — a curatable vocabulary with an embedded floor.
Hierarchical when the taxonomy is deep; coarse when unsure. The 75-industry cookbook's rule — report the fine category at ≥ 0.9, the parent otherwise — means the user always gets a useful answer and never a wrong-looking confident one.
Batch the catalogue nightly. The same questions over every existing row (re-categorise, find duplicates, check allergen terms, re-match orphan lines) is a map-reduce that costs pennies per tenant, and it is how the backlog RESEARCH §17 describes ("one-at-a-time pack reading does not clear a backlog") gets cleared for the categorisation half.
Log the answer beside what the user chose, from day one. That is the eval. Thresholds are read off the accuracy-per-confidence-band table after a fortnight, not guessed.
6. Money — what it costs, what it could save
All at list prices on 22 Sep 2026: Jev $0.042 per million input tokens, output free [verified]; Anthropic Haiku 4.5 $1 / $5, Sonnet 5 $2 / $10, Opus 5 $5 / $25 per million in / out [verified — the current Anthropic price list]. USD_TO_GBP = 0.79 as the functions use [repo].
Decision
Tokens per call (estimate)
Jev
The same decision on Haiku
On Sonnet
A. one receipt, 25 lines, 3 questions each, 6 candidates a line
~5,500 in, 0 out
$0.00023 (£0.0002)
~5,500 in + ~500 out ≈ $0.008 (£0.006)
≈ $0.016
B. one pack, 3 candidates, 1 Score + 3 Nouls
~600
$0.000025
≈ $0.001
≈ $0.002
E. one ProvenBot turn, 3 questions
~400
$0.000017
≈ $0.0007
≈ $0.0015
G. one declaration, 14 Nouls
~700
$0.00003
≈ $0.0013
—
G. nightly sweep, 300 declarations
~210,000
$0.009 per tenant
≈ $0.40
—
H. one feedback item, 4 questions
~500
$0.00002
≈ $0.0009
—
Read against what the app already spends [repo]: a receipt read on Haiku is of the order of £0.008 (the figure two failed Amazon-PDF reads both hit, #2147); a pack read on Opus is £0.0161 mean. So A adds about 2% to a receipt's AI cost; doing A on Haiku instead would roughly double it, and doing it on Sonnet would treble it. That is the whole cost argument for Jev over "just ask Claude": not that it saves money on today's four calls — it barely touches them — but that it makes the next fourteen decisions affordable to automate at all, at a price the £10-a-business daily ceiling would never notice.
The one direct saving is E. A ProvenBot turn carries a ~16k-token corpus, cached (read at 10%). A turn misrouted to Opus at effort high costs roughly the Opus cache read (~$0.008) plus thinking-and-answer output that can run to several thousand tokens at $25/M; the same turn on Sonnet at effort low is ~$0.003 cache read plus a few hundred tokens at $10/M. Of the order of 5 to 10 US cents per misrouted turn [inference]. Whether that is £5 a month or £500 depends on turn volume and the misroute rate, neither of which is measured — the anthropic-cost job (daily, by workspace) and api_usage_log by slug are exactly the instruments to measure it with, before and after. Side finding while reading the price constants [repo]: provenbot prices Sonnet at $3 / $15 and admin-api prices Opus at $15 / $75, both older generations' rates; the anthropic-cost job reconciles against Anthropic's own cost report, so this is a display drift in the estimate column, not a billing one — worth a one-line issue.
The larger saving is not in the AI bill. It is the onboarding minutes RESEARCH §17 measured as the reason a backlog does not get cleared, and the "which category?" click on every receipt line for a business that has 200 lines a week. Those are conversion and retention numbers, not token numbers, and the public opening (postponed, no date) is when they start to be measured.
7. Where it does not fit — recorded so nobody rediscovers it
Reading photos and PDFs — parse-pack, parse-receipt, parse-recipe on images. Jev is text-only [verified]. Opus stays on packs (the rule in the source and RESEARCH §26); Haiku stays on receipts and recipes. This is presumably why the first pass found "relatively few use cases": if the question was which of our LLM calls could Jev replace, the honest answer is none.
Writing anything — ProvenBot's answers, the recall explanation sentence, the admin drafting aids, release notes, replies. No generation [verified].
The label engine, the allergen decision, mandatory statements as decisions — derived, deterministic, human-confirmed (#15, #16, the labelling design doc). Jev is at most a checker feeding a queue (G).
Anything with arithmetic or dates — costing, VAT maths, use-by and best-before, shelf life, stock moves, the VAT-registration nudge, journey-watch statistics [verified — jaggedness 2, 3].
The data-quality Health check — rules over structured state; nothing to judge.
HACCP hazard types and critical limits — the app deliberately never suggests a limit (haccp.js:315); the user's own hazard analysis is the due-diligence record. A suggested hazard type is technically trivial and stays out on design grounds until Dave says otherwise.
A safety gate on ProvenBot tool calls — every writing tool is already propose_* behind a confirm card; the LangChain-style AutoMode gate solves a problem this app designed away.
Replacing Metricool, the FSA feed detection, or any ops routine's judgement — the routines are LLM sessions with connectors; Jev has none.
8. The costs of saying yes — line items, not discoveries mid-PR
A second product-runtime AI vendor, three weeks after deciding to have one. Decision 1939 (16 Sep 2026) retired the speech-to-text and social-scraping vendors "to keep things cleaner going into GA" and counts "product-runtime AI vendors 2 → 1 (Anthropic)". Jev makes it 2 again. That is Dave's call, and it is a different question from whether Jev is useful [repo].
A new sub-processor, with the notice before the first byte. TypeSafe AI, Inc. is US-hosted ("The Services are hosted in the United States") and will not train on inputs [verified — privacy policy]; the DPA offers EU SCCs Module 2 and the UK Addendum, names no retention period ("as long as necessary") and points to a sub-processor list behind a trust-centre login; zero data retention is enterprise-only [verified — DPA, docs.typesafe.ai/legal]. Our own Sub-processor List clause 3.2 promises an in-app notice before a new sub-processor begins processing Customer Personal Data, with a chance to object; privacy policy 7.1 carries the table; the D-06 disclosure test (ainotice660.js §5) pins that Settings → Data & privacy names the AI vendors. A receipt line is Customer Personal Data (it can carry a customer's name and address). So even a shadow pilot needs the row and the notice first, or runs on our own tenant's data only. The Cloudflare route (typesafe/jev on Workers AI) is a pass-through to TypeSafe's infrastructure [verified — Cloudflare docs], so it changes the billing and the secret, not the sub-processor.
A credential and a slug. Direct: TYPESAFE_API_KEY as a Supabase function secret, a credential-descriptions.yml entry before the first Deno.env.get() (the inventory guard fails otherwise), and api_usage_log rows under a system-one slug — kept out of AI_SLUGS so a Jev outage or a runaway loop can never spend the Anthropic ceiling, with its own tiny daily cap. Via Cloudflare: a Workers-AI-scoped token in the same place; the existing CLOUDFLARE_API_TOKEN is an Actions secret with Workers-edit scope and must not be reused.
One edge function, not one per use.system-one/index.ts: denyReason() like the others, the question sets versioned in the repo under named uses (receipt_lines, pack_match, provenbot_intent, …) so the client sends state and a use name, never questions — prompts stay reviewable, testable and pinned to jev-1.13.0. State is built server-side from the tenant's rows where possible, so a raw description never carries more than it needs.
A shadow log and the eval. A small table (migration, staging first, the pair committed before applying — #4) holding (use, business_id, record ref, answers json, confidence, what_the_user_chose, chosen_at). Two weeks of it is the accuracy-per-band table that sets the thresholds. pack_parse_events already gives B its ground truth.
Degrade to today, always. Every call has today's heuristic as its fallback and a 2-second timeout; a missing key means the feature is simply off. The vendor is a week old publicly with $40M and a 1,200 requests-a-minute limit that "adjusts with demand" [verified]; the open reproduction on a 4B-parameter model (0.845 modal agreement against Jev's 0.883 [reported — firecrawl.dev]) means the technique outlives the vendor, and the vendor's own critics say the decomposition is worth adopting on any model [reported — pearpages.com].
Untrusted state. Receipts, scraped pages and pack text are adversarial inputs (jaggedness 6). A Jev answer only ever pre-fills a field the user confirms or raises a finding for a human; it never commits a row and never bypasses a confirm card.
Copy. The user guide's receipt, pack and recipe chapters say what is now pre-filled and how to change it; a changeset per landing; the three-copy AI disclosure gains a sentence on "how lines are sorted".
9. The pilot, and the three calls that are Dave's
Recommended first pilot: A (receipt lines), shadow mode, then B. It has the highest click count, ground truth for free, a contained failure (a wrong category is corrected on the review screen the user already sees) and no compliance exposure. Two weeks of shadow answers, a threshold table, then the pre-fill switched on at ≥ 0.9. E (the router) second, because it is the only one with a direct saving and it can only fail safe.
Decided the same day (Dave, 22 Sep 2026).
Vendor: yes — the 1 October legal/publish gate was lifted 24 Sep 2026 (#2448); Deploy site no longer waits for 1 Oct for the TypeSafe notice. No customer row until that notice is live.
Route and notice: through Cloudflare (typesafe/jev on Workers AI, 32k context, no new vendor account); it is still a new sub-processor, so the clause-3.2 in-app notice goes out before any customer row is sent, shadow mode included.
First pilot:recipe lines (§4-C) in shadow mode — not receipt lines as recommended above; receipt lines, the pack picker and the router follow as their own children.
Filed: epic #2228 → #2229 foundation (#2234 sub-processor row, notice and AI disclosure; #2235 Workers AI token, secrets and credential inventory; #2236 the system-one function and metering; #2237 the shadow table and the readout) → #2230 recipe lines (#2238 shadow, #2239 switch-on), #2231 receipt lines (#2240, #2241), #2232 pack picker and duplicates (#2242, #2243), #2233 ProvenBot router (#2244, #2245). The allergen-scan second opinion (§4-G) is not filed: design first.
Sources
Fetched 22 Sep 2026 unless stated. TypeSafe: typesafe.ai (home), the launch post typesafe.ai/blog/introducing-system-one-models-and-jev, docs.typesafe.ai — introduction, quickstart, System One, State, Primitives (Choice, Score, Noul), Confidence, How to build with System One, Example use cases, Patterns (speculative fan-out, confidence-gated routing, composite scoring, intent routing), Models, API reference, Legal, Jev 1.13 jaggedness, and the cookbooks Parallel questions, Hierarchical classification, Knowledge graph entity alignment, SDE cascade, Pre-parsed value extraction, Classification using confidence, Re-ranking, Guardrails for LLMs, Classifying RAG passages; typesafe.ai/legal/privacy-policy, typesafe.ai/legal/data-processing; developers.cloudflare.com/ai/models/typesafe/jev/. The video: YouTube oEmbed for tYugqJ9YytQ (title, channel), the Acast episode page for AI News & Strategy Daily with Nate B. Jones, 21 Sep 2026 (show notes), natesnewsletter.substack.com (the companion guide's headline; body paywalled). Independent: DataCamp Jev: TypeSafe's System One Model Explained; DEV Community (Valyu) How to Use Jev; LangChain Building a harness with Jev; Tyler Folkman I ran 60 AI agents for a full day on Jev; Latent Space AINews on Jev; firecrawl.dev What is Jev (citing Every's tests and the openjev reproduction); pearpages.com Jev, sorted: what is fact and what is still a claim; flaviocopes.com A deep dive into Jev. ProvenBatch: the files and lines cited inline, all on main at the head of 22 Sep 2026; outputs/RESEARCH.md §17, §26; decision 1939 (issue #1939 / epic #1942); outputs/docs/ai-limits-console-design.md; outputs/tests/ainotice660.js; outputs/provenbatch-site/src/pages/sub-processors.md; Anthropic's current price list via the claude-api skill's cached table (24 Jun 2026).
Capability
Last updated 19 Aug 2026
Insights 2.0 — analytics research, competitor analysis and BI design
Researched 19 Aug 2026. Companion mockups: outputs/docs/insights-mockups/desktop.html and mobile.html. Backlog: epic #571 (Insights 2.0) carries §6's phasing; #572 is the phone-reachability bug from §1.
The brief (Dave, 19 Aug 2026): level Insights up into something the top tier buys into — deep research into what analytics would be valuable across all of the app's features (business success and compliance), the best way to display them so it feels like a real business-intelligence tool, flexible enough to show what users really want to see, covering at least what competitors offer. Desktop-first but genuinely usable on a phone.
1. Where Insights stands today
Insights became its own Pro-gated top-level tab in v0.205.0 (#548, 17 Aug 2026) on the stated premise "it will gain enhancements over time" — but no enhancement issues were ever filed. What ships today (from #35, v0.141.0):
Three sparklines — batch cost, revenue, profit — over a trailing 12 months or 4 tax years (insightsSectionHtml, index.html ~14809).
A per-product revenue / profit / margin table for the selected period.
A price-changes card (first-vs-last pack price per supplier product, sharing priceDrift()'s formula).
A one-line waste sentence (unsold units + logged stock waste, quantity only).
That is a good, honest v1 — but it is one card's worth of content holding down a whole tab that is now the newest Pro differentiator alongside Stock (#33) and Food safety (#36). Two immediate observations:
Insights is unreachable on a phone. The phone Menu (renderMore, index.html ~5096) hand-maintains its row list, and the Business group stops at Accounts — #548 added the sidebar button but no Menu row. A Pro subscriber on a phone cannot open the tab at all. Filed as #572 — it's a reachability bug ("a feature is only shipped where it can be reached", DESIGN-LANGUAGE), not part of this design.
The data is already there. Almost every metric in §2 below computes from tables the client already loads (cache.* / window._batches) using helpers that already exist (summaryForRange, batchFigures, recipeComputed, insightsPeriods). Insights 2.0 is overwhelmingly a rendering and product-design problem, not a data-engineering one. There is no server analytics layer and this design deliberately does not require one.
House rules every metric below already respects
These are test-enforced or Dave-ratified; the catalogue was filtered against them rather than flagging violations downstream:
Rule
Source
No composite scores — no business health index, no compliance score. Counts per class, or nothing.
Reuse the period spine — monthRange / taxYearRange / inRange, grown with a count parameter; never a second boundary implementation.
insights35.js asserts; #349
Partial data omits, never zeroes — a batch with no computable cost contributes nothing; est/adj/"can't work it out" flags carry downstream to every derived figure.
#204 discipline
No invented £ — waste stays physical quantity; no carbon series (dropped 9 Aug 2026, no defensible UK data source); nutrition deferred with #28.
#35 decisions
A second reading of a fact reuses the first's formula — e.g. price-per-pack is amount ÷ qty everywhere.
DESIGN-LANGUAGE
Compliance is never gated and never scored — attestations (4-weekly review) are never turned into "% compliant"; nothing on a compliance surface is greyed by plan.
D-05 hard rule; plan207.js
A count equals the filtered list it links to — every number drills to the exact rows that produced it, via tableFilters, same predicate.
#540 counts-first rule
Client-side aggregation over cache — no warehouse, no server rollups (cross-tenant work belongs to epic #551, out of scope here).
Architecture
2. Metric catalogue — what's valuable, per feature area
Everything below names its data source in today's schema and the helper that already computes it (or the nearest one to grow). Phase refers to §6. Metrics are grouped into the four domains the proposed UI uses: Money · Making · Selling · Compliance.
The strategy anchor for the mix (RESEARCH §2): the buying trigger is compliance fear, not efficiency — "prove you did it right", not "save time". So compliance analytics sit beside the money story as a peer, not behind it. That combination is also the competitive white space (§3): business-BI products have no compliance story and compliance products have no money story.
recipes.batch_time_mins as a labour proxy against product margin — "your highest-effort, lowest-margin products"
cache.recipes, insightsProductFigures
Only for recipes with a recorded time; never imputed
3
2.3 Selling
Metric
Definition
Data source / helper
Notes & constraints
Phase
Orders funnel
Count/value by status (New → Quoted → Confirmed → Ready → Collected; Closed) for the period
cache.orders, ORDER_STATUSES
Pure counts; the funnel is a table+bars, not a chart-library funnel
1
Quote → confirmed conversion
Share of quotes that confirm, and median days to decision
orders status history via created_at, quote date, status
If a timestamp isn't recorded, the order is omitted from the rate
2
Close reasons
Breakdown of why orders didn't proceed (declined / no response / cancelled / error)
orders.close reason
Answers "where am I losing work" — no competitor at this size offers it
1
On-time fulfilment
Collected on/before due_date vs late
orders.due_date, collected_at
Both dates required, else omitted
2
Average order value & order count trend
AOV per period; orders per period
incomeOrdersInRange
Reuses Accounts' income-date rule (incomeDateOf)
1
Customer value / repeat rate
Revenue per customer, new vs returning per period, time-since-last-order
cache.orders × cache.customers
Square's customer framing, from data already held
2
Product mix in orders
Units by product across order lines; rising/falling movers
cache.orderItems
The only reader orderItems would ever have had
2
Unpriced-order leakage
Collected orders with no price (the Accounts "heads up", trended)
orders where collected & price null
A count that drills to the fixable rows
1
2.4 Compliance
The domain competitors' compliance tools monetize hardest (§3.4), and ProvenBatch's data here is unusually rich: an append-only entry log with attests-to vs recorded-at, sign-offs, corrective actions, allergen verification history, and weekly health-check snapshots.
Metric
Definition
Data source / helper
Notes & constraints
Phase
Check completion rate
Recorded ÷ due per day/week, honouring trading pattern and exceptions
THE "prove you did it right" number; already computed piecemeal on the Allergens tab, never trended
1
Verification currency trend
Declarations past their review window over time
Same + compliance_log timestamps
2
Second-opinion disagreement rate
Scans where the second opinion flagged something, per period
Second-opinion results (compliance_log)
Framed as "caught before it shipped" — a positive
3
Health-check counts by grade
Critical / Major / Minor trend (exists on the Health check tab)
cache.dqWeeklySnapshots, dqTrendHtml
Cross-linked from Insights, not duplicated — one formula, one home; Insights shows the current counts + a link
1
EHO-readiness checklist
Named counts: unverified allergens on sold products; declarations past review; open corrective actions; days since last 4-weekly review; missing sign-offs this month
All of the above
A checklist of counts, each drilling to its fix — deliberately not a score (§7)
1
HACCP review currency
Plans with no review inside their cycle
cache.haccpPlans, haccpPlanReviews
Count + names
3
2.5 Cross-cutting: "What changed" (the narrative layer)
The single biggest UX gap between today's Insights and a "real BI tool" feeling isn't more charts — it's explanation. Shopify's Spring '26 release added exactly this ("Insights" annotations that say why a number moved). For ProvenBatch it is cheap, because the causes are in the same cache as the effects:
Butter (Wyke Farms 250g) rose 9% on 12 Aug — margin on 4 products fell, Brownie box now 31% (was 36%).
3 opening checks missed last week, all Tuesdays.
2 declarations passed their 12-month review — Verify.
Orders from repeat customers: 64% this month (was 51%).
Each line is derived from one already-computed metric, names its records, and links to the filtered surface that fixes or explains it (house rule). This is a rules-based, deterministic list — not an LLM feature, no API cost, no #204 risk. Phase 1 ships it with 4–6 rule types; more rules accrete over time.
3. Competitor analysis
All checked 19 Aug 2026 via vendor sites/support centres unless noted; marketing-level claims flagged. Competitor research is maintained in the master competitor sheet.
3.1 Labelling / compliance SaaS
Product
Analytics offering
Tier gating
Nutritics ("Insights" module)
Build-and-customise dashboards across modules: menu profitability, waste, carbon; 50+ pre-configured layouts; Usage Reports — find all recipes containing a given ingredient/allergen
Premium only; fully-custom dashboards a further upgrade
FoodDocs
AI Reports: heatmaps + weekly graphs of task completion per location/team/task; one-click CSV/XLSX historical export "for auditors in seconds"
Export at Professional ($299/mo); AI Reports at Enterprise
Safe Food Pro
Management dashboard (current/historic compliance state); per-form completion rates with drill-down to individual forms; one-click compiled audit document; auditor dashboard access; group (multi-venue) rollups + AI summaries
Not visibly gated (all-inclusive pricing)
Navitas Safety
Traffic-light dashboard (red/amber/green) fed by sensor pods; customisable reports by supplier/staff/checklist/appliance; scheduled report delivery (SFTP)
Enterprise sales, pricing not public
Squizify
Real-time dashboards, site comparisons, historical trends, standard + bespoke dashboards, team scorecards, automated audit-ready reports
Not public
LabelLogic Live (Planglow)
Downloadable reports from print lists: product usage per list, quantities printed, pricing, ingredients/allergens — print-volume-as-production-telemetry
None — £15–20/mo all-inclusive
Erudus
Query Builder over the shared product-data pool; PDF/CSV report generation. No dashboards/KPIs
Recipe cost per package/batch incl. labour/overhead; cost-over-time view with target margins; nutrient-driver breakdown
Tiers gate counts, not analytics
Genesis Foods (Trustwell)
Data distribution via API to external reporting systems — no BI of its own
Enterprise
(Corrections to our prior competitor set: Bakord is an Indian white-label grocery-delivery platform, not a bakery tool — dropped. Menu Guide (menuguide.pro) is a standalone QR-menu product, not a Chomp feature, with no analytics.)
3.2 Bakery / food-maker management
Product
Analytics offering
Notes
Cybake (RedBlack, UK)
The BI benchmark at this size: Power BI built into Cybake 4; Retail BI (EPOS sales analysis, sales-based ordering, waste trends); Wholesale BI (profit by customer and by product, margin comparison over time, purchasing KPIs, supplier price/performance)
BI is a paid add-on module
Craftybase
COGS from actual purchase prices; P&L, expenditure/revenue, inventory valuation & turnover, product performance, high-revenue-customer reports, per-order profitability, pricing guidance auto-updating when material prices change
$49–349/mo; advanced ops features push tiers up
Streamline (Mountain Stream, UK)
User-built pivot tables, charts and tailored reports; gross profit per item AND per customer; custom attributes as report dimensions
Wholesale ERP; pricing not public
FlexiBake
BI module: forecast-vs-actual, sales analysis
ERP price point
BakeSmart
Automated daily production reports; "real-time analytics" (unspecified)
OrderNova
Date-range order/production reports with tallying/grouping, print/export
The low end has nothing — our Starter/Standard users come from here
FoodCore (from our existing analysis)
Supplier price history, stock alerts, production runs; allergen matrix + food safety gated to its £65 top tier
The direct UK competitor; £25/£40/£65
3.3 What small-business owners are trained to expect (Square / Shopify / Xero)
Square: free sales-summary dashboard with this-vs-last framing on every card, item sales, new-vs-returning customers, visit frequency; ~10 stock reports + saveable custom reports (report blocks composed into one view); a daily sales summary email; deeper vertical reports paid.
Shopify: overview dashboard + reports library free everywhere; depth is the tier lever (cohorts, attribution, saving custom reports = Advanced/Plus). Spring '26 added "Insights", annotations and metric targets — the dashboard now explains why numbers changed. That is where the UX bar is heading.
Xero: free business snapshot (income/expense trends, profitability) + short-term cash flow; Analytics Plus (top-tier/add-on) extends horizon, prediction and snapshot customisation.
An owner who uses Square by day reads "vs last month" delta chips and taps into the number behind every card. Those two behaviours are the expectation floor.
3.4 Synthesis
Table stakes (users will assume these exist):
KPI overview with period-over-period deltas
Product/sales performance ranking
Cost & margin per product/recipe over time
Compliance completion rates + missed/overdue
One-click audit-ready export (we largely have this: EHO pack, accountant pack, allergen CSV — Insights should surface them, not rebuild them)
Drill-down from any summary to its records
CSV export of any table
Differentiators nobody in the space owns (ranked by fit to our data):
Money + compliance in one insights surface. Business-BI products (Cybake, Craftybase, Square) have zero compliance; compliance products (FoodDocs, Navitas) have zero margin. For a PPDS-driven buyer, one tab that answers "am I making money and can I prove I'm safe" is unowned.
Ingredient blast radius — price rise → affected recipes/products/margins (only Nutritics has the query form, gated Premium).
Narrative "what changed" insights — deterministic explanations, the Shopify Spring-'26 pattern at 1/100th the price point.
Weekly owner digest email — almost absent across the whole space (Square's daily summary is the closest).
Print/label telemetry as production trend — we have real batch + label data, no POS required.
How the space gates analytics — analytics is the top-tier carrot almost everywhere: Nutritics Premium, FoodDocs $299+/Enterprise, Cybake paid module, Xero Analytics Plus, Shopify Advanced. The consistent gating levers: dashboard customisation, export depth, history horizon, scheduling/digests, AI summaries, multi-site rollups. Pro placement of Insights is validated — and the levers suggest what belongs in Pro forever vs what could ever tease downward (§5).
4. The BI experience — how it should look and feel
Design goal: a real BI tool, not a page of charts — while staying inside the house design language (tokens, cards, pills, drawers, sentence case, hand-rolled SVG). The mockups render everything in this section.
4.1 Architecture: Overview + four domains
One tab, five views via a segmented control (the same .seg pattern as Month/Tax year today — no new nav):
Insights
[ Overview | Money | Making | Selling | Compliance ] [ Month | Tax year ] ‹ Aug 2026 ›
Overview (landing) — the owner's one-glance page:
KPI row — revenue, profit, margin %, orders, compliance attention count — each with a vs-prior delta chip (▲/▼, green/red by desirability not sign). KPI tiles on desktop, pills on phone (existing .kpi-row behaviour).
"What changed this period" — the narrative list (§2.5). Each row: plain-English sentence, severity pill, tap → the filtered fixing/explaining surface.
Pinned cards — the user's own selection (see flexibility, §4.3). Default pin set: revenue trend, margin by product, check completion, verification coverage.
Money / Making / Selling / Compliance — each a column of cards from §2's catalogue: a headline figure + delta, a chart, and a supporting table that drills down. Cards render in a .tcards-style responsive grid (one column phone, auto-fill ≥900px).
Why not user-composable dashboards à la Nutritics? Because a blank canvas is the "wallpaper" failure for a solo baker — curated cards with pinning gives 90% of the perceived flexibility ("shows what I really want to see") at 10% of the build and none of the empty-state risk. Pinning is the industry's customisation lever, delivered in house idiom.
4.2 Display patterns (the chart family)
All hand-rolled inline SVG, grown from insightsSparkline's idiom (viewBox, width:100%, token strokes, aria-hidden with the accessible number adjacent):
Pattern
Use
Anatomy
Sparkline+
All trends
Existing polyline + a soft area fill, last-point dot, min/max gridline pair with end labels (two lines, not a grid), period labels at the ends only
Delta chip
Every KPI and card headline
▲ 12% vs Jul pill; green/red by desirability (cost ↓ is green); — when prior period unknown
Bar row
Rankings and funnels (margin by product, orders by status, close reasons, spend by category)
Label · value · proportional bar in a table row — reads as a table, scans as a chart
Dot strip
Temperature readings vs target band
Band as a soft rect, readings as dots, out-of-range in --red
Heat calendar
Check completion
The existing .fsmonth pattern, reused not rebuilt
Donut — rejected
Proportions render as bar rows; a donut hides small categories and needs a legend
Interaction: tap the card → drill. A trend card opens its per-period table; a ranking row opens the record's editor or the filtered list (openChainList / tableFilters patterns). No hover tooltips as the primary interface — they don't exist on touch, and the design language already refuses per-point tooltips. Values a finger needs sit in the adjacent table.
4.3 Flexibility — "show what users really want to see"
Pin to Overview — a ☆ on every card header; pinned set + order stored per user. This is the whole customisation model, deliberately.
Per-card period override — cards default to the page period; a card can be flipped to 12-month view without leaving the page (the existing seg control, miniaturised).
Every table exports CSV — the existing downloadCsvText plumbing; one ⬇ per table.
Saved context, not saved reports — the page remembers its last domain + period per user. (Full saved-report composition is a rejected v1 feature; revisit only on user demand — Shopify gates exactly this at Advanced, so it's a future lever, not a launch need.)
Weekly digest email (phase 3) — the owner's Monday email: last week's KPI row + top 3 "what changed" lines. Rides the existing lifecycle_emails machinery. Off by default, one toggle.
4.4 Mobile
Same render, two shapes (.only-phone/.only-wide where markup differs; CSS elsewhere):
KPI tiles → the existing KPI pill compaction.
Domain seg control scrolls horizontally; cards stack single-column; tables ride .tw horizontal scroll.
"What changed" is better on the phone than the desktop — it's the glanceable, actionable layer; it leads the page at both widths.
Prerequisite: the phone-Menu reachability fix (own issue), which is a two-line change to renderMore and independent of everything else.
4.5 Honesty furniture
Every card inherits the #204 vocabulary: est/adj pills on figures with estimated inputs; "can't work it out — needs a yield" info rows instead of bare dashes; totals that name how many rows they omit ("from 14 of 16 batches — 2 uncosted"). This is a feature, not a caveat: no competitor states its own data quality inline, and for a compliance-fear buyer, a tool that visibly refuses to make numbers up earns the trust the whole product sells on.
5. Tier positioning
Insights stays Pro (per #548, validated by §3.4 — analytics is the industry's top-tier carrot). Pro = Stock + Food safety + Insights + 10 seats + priority support; Insights 2.0 makes that a story: "Pro runs your business by the numbers."
The lock card gets a preview. Today's featureLockCard("insights") is a sentence and a button. Show the real Overview beneath it rendered from the business's own data but blurred/static beneath the lock card — competitors gate the depth, we can gate the whole while still showing what's behind the door. Complies with #342 (no nagging, one card, real prices from PRICES_GBP) and never gates anything compliance-critical: every compliance surface (Allergens tab, Food safety diary, Health check) keeps its own ungated home — Insights only re-reads them. If the blur reads as taunting in UAT, fall back to a static sample-data screenshot card.
Compliance metrics duplicate nothing and gate nothing. The EHO-readiness counts all exist ungated on their home tabs; Insights aggregates the reading. A Starter business loses no safety capability by not having Insights — D-05's line holds.
Recommendation (separate small issue, not this epic): Settings' plan rows list only product/seat caps — a buyer comparing tiers can't see that Pro includes Insights/Stock/Food safety. One line per plan row fixes the top-tier story at the point of purchase.
6. Phased roadmap
Sized so no phase ships wallpaper (#35's lesson) — every card in a phase has honest data behind it on day one.
Phase 1 — the Overview + trend depth (M). Segmented domains; KPI row with delta chips; "what changed" v1 (price rises → margin impact naming products; missed checks; verification due; unpriced collected orders); sparkline+ upgrades; Money/Making cards from shipped data (trends, price changes, throughput, sell-through, waste trend); Selling v1 (funnel, close reasons, AOV); Compliance v1 (check completion trend, corrective-action aging, verification coverage, EHO-readiness counts, health-check cross-link); CSV per table; phone reachability fix lands with or before this. Phase 2 — the differentiators (M/L). Ingredient blast radius; batch cost variance; margin floor + ranking; quote conversion, on-time, customer value/repeat, product mix; temperature dot strips; late/skip rates; sign-off streaks; pin-to-Overview. Phase 3 — the retention layer (M). Weekly digest email; VAT runway; stock turns/days-of-cover; batch-time vs margin; second-opinion disagreement; HACCP currency; per-card period override if not landed earlier.
Each phase is a separate issue under the epic; each claims per the claiming rule before build.
7. Explicitly rejected (do not re-propose)
Rejected
Why
Composite scores — business health index, compliance score, audit-readiness %
House rule, dq204.js-asserted. §3 shows a numeric "compliance score" is open ground competitors haven't taken — we deliberately leave it open: a single figure summing margin and missed fridge checks cannot decompose to an action, and a scored attestation is a fiction. The EHO-readiness checklist of counts delivers the same reassurance honestly.
Carbon series
Dropped 9 Aug 2026 — no defensible UK ingredient-level data source. Nutritics sells this; we don't follow (their number rests on category factors — the #204 "confident number on a guess" class).
£ figure on waste
No honest per-unit cost basis at the moment of loss. Quantity only.
Charting library
House rule; the §4.2 family covers every §2 metric with hand-rolled SVG.
Blank-canvas dashboard builder
Wallpaper risk for a solo operator; pinning delivers the felt flexibility. Revisit only on real user demand.
LLM-generated insight text
"What changed" is deterministic rules over cache data — no API cost, reproducible, no hallucination surface on a trust product.
Cross-tenant benchmarks ("businesses like yours")
Genuinely differentiating, but parked with L7 price benchmarks (system-wide-learning-plan §7): four hard gates incl. ~20 active tenants and double opt-in. Post-beta at the earliest.
8. Loose ends surfaced by this research (not this epic's work)
Migration 548 (Insights → Pro) is applied to staging only — production apply must ride the next promote (flagged on #548's thread at the time; still open as of 19 Aug).
Four docs still say Insights = Standard: go-to-market-decisions.md (D-05 table ×2), feature-entitlements-431.md §2 D1, provenbatch-website-content-foundation.md pricing table (~line 275). Sweep before any pricing-page/legal work reuses them.
Settings plan rows don't list feature entitlements (§5, own issue).
Mobile PWA research — where the app breaks on a phone, and why
1 Aug 2026 · researched against v0.92.0 (production) · screenshots in mobile-shots/, reproducible via outputs/mobile-harness/ (README there)
Dave's report: "a number of pages don't fit well onto a mobile screen — overlaps and pieces that get cut off." This document confirms that, finds the mechanisms (there are two, not twenty), maps every affected screen, and proposes fixes in three shippable batches.
How this was researched
The real app — v0.92.0's index.html, byte-identical bar three mechanical patches — was booted in Chromium at 390×844, 360×800 and 320×568 (iPhone 12–15, common Android, iPhone SE1), signed in as a seeded tenant with realistically-long data ("Bookers Cash & Carry (Aylesbury)", "Orange & Polenta Celebration Cake (gluten free)"). Every screen and major drawer was visited through the app's own navigation, screenshotted, and measured by an in-page audit that flags (a) elements crossing the right viewport edge with no scrollable ancestor and (b) innerWidth exceeding the device width. This goes beyond the July UX evaluation (#303), which reconstructed individual screens statically: here the whole app runs, so cross-screen mechanisms (a stretched page displacing a drawer) are visible. No real environment or data was involved — outputs/mobile-harness/README.md explains the stub.
Scope note: everything below is layout at phone widths. The #303 epic's findings (navigation, capture flow, confidence routing) are separate and mostly shipped in v0.86.0.
The two mechanisms
Nearly every defect found reduces to one of two shared causes. Fixing the causes fixes screens this research never looked at.
Mechanism A — tables are silently amputated by .card{overflow:hidden}
The app has 36 <table> emission sites and a shared table style with no mobile strategy. Only a handful (the drafts list, the expenses table, the price-compare matrix, Help's .tw) are wrapped in an overflow-x:auto container. Everywhere else the table sits directly in a .card — and .card clips (overflow:hidden, index.html ~line 197). At phone widths the table is wider than the card, so the rightmost columns simply do not exist for the user: no scroll, no affordance, no hint. And because table layout squeezes the columns that remain, long names wrap one word per line.
What gets amputated is not decoration — it is the money and the actions:
Screen
Columns the phone user cannot see or reach
Batches (the core operational table)
Made, Batch cost, Sold, Profit
Products
Sale price (cut mid-glyph), Margin, Margin %, Status
The Allergens case deserves its own line: the Second opinion re-reader is a compliance surface, and its resolution actions are the part that is off-screen. A phone user can see that the app disagrees with a stored allergen level but cannot act on it.
Mechanism B — one overflowing element stretches the page, and everything fixed goes with it
Three screens contain an element that escapes its card and stretches the document itself past the device width: Recipes (530px), Products (451px), Accounts (549px) — the culprits are their filter/toolbar rows (search + select + primary button in a no-wrap flex row) and, on Accounts, the two-up card grid whose inner tables push it wide.
When that happens, mobile Chrome does what mobile browsers do with overflowing pages: it treats the page as wider than the phone. Text shrinks, and — this is the part that produces Dave's "overlaps and pieces cut off" — every position:fixed element now anchors to the stretched page, not the phone screen. Measured directly: on Accounts, innerWidth becomes 549 on a 390px device, and the drawer (right:0; width:100vw) opens at x=159..549 — 40% of the New expense form hangs off the right edge of the phone, including the Attach button, the item picker and the TOTAL £ field. The same displacement hits the Edit recipe drawer opened from Recipes, and the tab bar/Capture FAB render off-centre while it persists.
The control experiment proves the drawers themselves are fine: the same drawer machinery opened from an honest page (Ingredients) fits the phone exactly — full width, two-column grids collapsed, chain rail wrapping properly.
Also visible in that shot: the sticky action row's last button ("Close") is clipped — the four-button footer row does not wrap (F7 below).
The immediate trigger on each stretch screen: the primary action gets pushed off-screen — "+ New product" is clipped to a sliver on Products (see above) and Accounts loses its second toolbar button (the #315-restored "📷 Scan receipt") behind the right edge:
The M1–M3 investment shows. Today, Capture, the nav drawer, Look up, More, Settings, Inbox, Shopping, Quality, Sessions, Help and the receipt-review flow are genuinely good on a phone — cards stack, forms collapse, type stays readable, the tab bar behaves, bottom padding clears it when scrolled. The problem is confined to the desktop-era table screens and what they do to the rest of the page.
F1 · Tables amputated (Mechanism A). All screens in the table above. Severity: high — silent data loss on the operational core (Batches profit figures) and a compliance surface (Allergens actions).
F2 · Page stretch displaces fixed UI (Mechanism B). Recipes, Products, Accounts + every drawer opened from them. Severity: high — forms become part-unreachable; this is the single biggest source of "overlap/cut off" reports.
F3 · Toolbar/filter rows never wrap. Search + select + button rows on Recipes, Products, Accounts (and the Batches search input truncating its own placeholder). Severity: high on the two screens where the clipped control is the primary action.
F4 · Receipt-review photo pane overlaps itself. In the (otherwise excellent) Review receipt drawer, the photo pane's title/pills collide with the ✂️ Adjust button and the caption is crushed to a one-word-per-line column — the buttons row (Adjust, + Add a page, Discard) does not fit 390px and does not wrap. Severity: medium (cosmetic-to-annoying; everything still tappable).
F5 · Long names wrap pathologically inside crushed columns. "Orange & Polenta Celebration Cake (gluten free)" renders as seven lines of one-to-two words in Recipes/Batches/Products. A consequence of F1's column squeeze, worth its own acceptance criterion (a name column keeps a usable minimum width; dates shorten — 2026-07-25 → 25 Jul — before names give way).
F6 · ISO dates burn column width.2026-07-25 wraps to two lines in every date column. Every date cell should render the app's short form on phones.
F7 · Drawer sticky footers clip their last button. Four-action rows (Save / Not an ingredient… / Delete / Close) run off the right edge instead of wrapping. Severity: low-medium (Escape and ✕ still close), but it reads as broken.
F8 · Search placeholders overrun their inputs. "Search recipe, session or l…", "Search supplier or referen…" — cosmetic, but the truncation is visible in every screenshot. Shorter placeholders on mobile ("Search…") cost nothing.
F9 · 320px works exactly as badly as 390px, no worse. The failures are the same class, just tighter (topbar name truncates to "Hedgerow P…", which is #314's intended sacrifice order). No SE-specific defects found — fixing 390 fixes 320.
Proposed fix plan — three batches
Batch 1 — stop the bleeding (one version, mechanical, low-risk).
Wrap every data table in the existing scroll-wrapper pattern (overflow-x:auto container / .mx-scroll / Help's .tw — one class, used everywhere). This makes every amputated column reachable immediately and — because the table then scrolls inside its container — removes the page-stretch triggers, which fixes the displaced drawers for free.
flex-wrap: wrap on the toolbar/filter rows (Recipes, Products, Accounts, Batches search) so primary actions drop to a second row instead of leaving the screen; drawer sticky footers and the receipt-pane button row get the same treatment (F4, F7).
Short date format in table cells at phone widths (F6); shorter mobile placeholders (F8).
Regression guard, house style: a new suite (e.g. tests/mtables.js) asserting statically that every <table emission site sits inside a scroll wrapper — the same grep-the-source pattern sw169.js uses, so the rule cannot quietly rot. The harness (outputs/mobile-harness/) is the measured half for manual re-verification; JSDOM cannot do layout, so the static rule is what CI enforces.
Batch 2 — a real mobile treatment for the four workhorse tables (per-screen, design work). Sideways scroll is a floor, not a ceiling: a column you must discover by scrolling is a column most users never see. At ≤640px, render Batches, Recipes, Products and Orders as stacked card-rows (name + status pill on the first line, the two numbers that matter on the second, chevron to open — the pattern Today's cards already use), per DESIGN-LANGUAGE. The Allergens second-opinion table should become stacked rows too — its actions must be on-screen, not scrolled to (compliance surface). Desktop keeps the tables. This is the same shape as #303's sub-issues: one issue per screen, one version per group.
Batch 3 — polish. Accounts' two-up card grid stacks to one column at phone widths; audit the remaining minor tables (ingredient-detail chain, run detail) for the card-row treatment; re-run the harness and archive the after screenshots beside mobile-shots/.
Suggested backlog shape: a small epic ("Mobile fit & finish — the table screens"), Batch 1 as one issue (it is one coherent change + its test), Batch 2 as four-to-five sub-issues, Batch 3 as one. Batch 1 is safe to ship immediately and makes the app honest on a phone; Batch 2 is where it becomes good on a phone.
Appendix — per-screen result (audit, all three widths)
clip = elements beyond the right edge with no scrollable ancestor (Mechanism A) · stretch = page wider than the device, viewport zoomed out (Mechanism B) · identical results at 390/360/320 unless noted.
Screen
Result
Today, Look up, More, nav drawer, Capture sheet
✅ clean
Shopping, Quality, Sessions, Inbox, Settings, Help
✅ clean
Suppliers
✅ clean (its list fits)
Batch (+ batch detail, run detail)
clip — costs/profit/actions amputated
Orders
clip — status/price amputated
Ingredients (+ detail chain)
clip — chain columns amputated
Recipes
stretch to 530px + clip — costs amputated
Products
stretch to 451px + clip — margin/status amputated, + New product clipped
Allergens (second opinion + matrix)
clip — resolution actions off-screen
Customers
clip — contact/orders amputated
Accounts
stretch to 549px — toolbar button lost, card grid crushed
Review receipt drawer
fits; photo-pane overlap (F4)
Edit drawers from honest pages
✅ fit (footer wrap aside, F7)
Any drawer from a stretched page
displaced off-screen (F2)
Method caveat: screenshots use seeded data sized like a real account a few months in. An account with longer supplier names or more columns-worth of data will be strictly worse; an emptier account hides most of this — which is why it survived on-device checks so far: the demo-sized accounts used in testing don't stretch the tables.
Shipped state (1 Aug 2026 — epic #327 complete, all three batches)
All three proposed batches above shipped as filed, plus one addition Batch 3 found: the Ingredients list (named in the appendix above as amputated) wasn't covered by any of Batch 2's four sub-issues — an oversight in the epic's own scope split — so Batch 3 gave it the same card-row treatment as Batches/Recipes/Products/Orders/Customers, and gave the Data quality screen's three-column What/Why/Fix rows (not flagged by the harness — text wraps rather than overflowing — but genuinely hard to read at phone widths) a stacked-card treatment too.
327a (Batch 1, v0.93.0): every <table> scroll-wrapped, toolbars/drawer footers/receipt pane wrap, short table dates, mtables.js guard.
327b–327e (Batch 2, v0.94.0): Allergens second-opinion, Batches (+run detail), Recipes, Products, Orders, Customers all become stacked card-rows at ≤640px, desktop unchanged.
327f (Batch 3, v0.95.0): Accounts' Income/Expenses grid stacks to one column; Ingredients (the missed screen) and Data quality's findings rows get card treatments too; harness re-run clean; after screenshots archived in mobile-shots/after/.
Re-run result: 110 of 111 screen×viewport states clean at 390/360/320px — zero clip, zero stretch, on every screen this document names. The one remaining state (help at 320px only) is a small, unrelated, pre-existing overflow on the Help screen's inbound-email <code> block (a generated address running ~26px past the narrowest supported viewport, already word-break: break-all) — not a table, not a toolbar, outside this epic's remit. Filed as its own small follow-up rather than fixed here.
Screen
Before
After
Batch (+ batch detail, run detail)
clip — costs/profit/actions amputated
✅ card-rows, after/batches.png
Orders
clip — status/price amputated
✅ card-rows
Ingredients (+ detail chain)
clip — chain columns amputated
✅ card-rows (Batch 3), after/ingredients.png
Recipes
stretch to 530px + clip
✅ card-rows, no stretch, after/recipes.png / after/recipes-320.png
✅ no stretch, grid stacks to 1 column (Batch 3), after/accounts.png
Review receipt drawer
photo-pane overlap (F4)
✅ header row wraps
Any drawer from a stretched page
displaced off-screen (F2)
✅ nothing stretches any more, after/drawer-recipe.png / after/drawer-expense-new.png
Data quality (What/Why/Fix rows)
not flagged, but crushed to single words
✅ stacked cards (Batch 3 addition)
Help (inbound-email code block, 320px only)
—
⚠️ still overflows ~26px; unrelated, not fixed
Epic #327 is closable on this evidence — every screen named in this research's original findings (F1–F9) is now clean or, in the one Help case, explicitly filed as a separate follow-up rather than silently left.
Capability
Last updated 3 Sep 2026
VAT-registered as a business setting — what would actually change (#1110)
Researched 3 Sep 2026, against main at v0.301.3. This is the research Dave asked for on 2 Sep 2026 ("'VAT registered' should be a setting in business settings, and VAT should be managed consistently across the app based on it — but needs research on what is needed where for VAT registered vs non businesses"). It is a research/decision record, not a build: no schema, no UI, no code changed by the issue that produced it.
Status, 3 Oct 2026: stage 0 shipped as #1146. Stage 1 is built as #1147 — Dave lifted the evidence gate for the farm-shop segment (epic #3055) and decided unit_price is VAT-INCLUSIVE. The build's decisions, and where it departs from §5 below, are in outputs/docs/decisions/1147-output-vat-pricing-basis.md. This record is otherwise left as the research it was.
Key takeaways
A vat_registered boolean on its own buys one thing: it stops the #354 threshold strip nagging a business that has already registered. Everything else people expect it to fix — a legal VAT invoice, correct margins, correct turnover — needs the output side of VAT, which the app does not model at all and which is where all the cost is.
The output side has a real, unnamed correctness defect waiting in it. For a VAT-registered business selling standard-rated output, unit_price silently contains 20% that is not the business's money, so Margin %, batch economics, Income and Net profit are all overstated, and SA103 turnover with them. No part of the app says whether a price is VAT-inclusive. This is worth more than the invoice question and nobody has written it down before.
The zero-rating wrinkle cuts the problem roughly in half, but not the way the issue assumed. Most cold bakery output is zero-rated, so for that business registration changes almost nothing on the output side. But "small food business" is not "cold cake maker": caterers, hot-food stalls and anyone selling confectionery are standard-rated on everything, and for them registration changes a great deal. The setting cannot assume zero-rating.
**Zero-rated businesses are more likely to register voluntarily than the £90k threshold suggests**, because they charge 0% and reclaim input VAT — a permanent repayment position. So vat_registered must not be derived from, or gated behind, the turnover nudge. They are independent facts.
Recommendation: stage it.business_settings.vat_registered alone now (one column, one consumer, cheap, fixes a live wrongness); the whole output-VAT layer — vat_number, per-product liability, the VAT-inclusive/exclusive decision, VAT invoices — held as awaiting evidence until a real business tells us they are registered and want to invoice from ProvenBatch. Nobody has.
Not blocked on #826. Ship UK-only, the same posture #354 and #901 already shipped on.
🔴 MTD submission stays permanently out of scope (#829, accounting-capability-research.md §4). Registration is exactly the moment someone will argue otherwise. Knowing a business is registered creates no obligation for us to file — see §7.
1. What the app knows about VAT today
Verified by reading the code, not from the issue's summary.
Where
What exists
Side
VAT_REG_THRESHOLD / VAT_NUDGE_AT (part-2.js:1461)
£90,000 and 0.8, as data in one place; vatNudgeHtml() renders an Accounts strip from trailing-12-month income (#354)
neither — a turnover fact
VAT_RATES = [20, 5, 0] (part-2.js:1466)
The UK purchase-rate vocabulary, sent to parse-receipt as a #279 contract (#901)
Sum of recorded vat_amount, shown only when businessTracksVat() finds any VAT data at all
input
Accountant export
Net / VAT / VAT rate % columns (#901)
input
orderDocHtml() — invoice and quote
🔴 No VAT line at all, by decision (#1045, Dave, 2 Sep 2026)
output
Everything else
Nothing. No registration flag, no VAT number, no per-product liability, no output rate, no tax point
—
The asymmetry is the whole story. The input side was built (#901). The output side was considered once, on #1045, and deliberately deferred. There is no flag in between.
1.1 The one thing the app says out loud today
outputs/USER-GUIDE.md:2037 — "ProvenBatch … does not know whether you're VAT registered, and does not submit anything to HMRC — and it isn't going to."
That sentence is accurate today and is the sentence #1110 would change. Note it is doing two jobs at once: stating a data-model fact ("does not know") and a permanent commitment ("isn't going to" submit). Only the first half is in play here. If a flag ships, that sentence must be split, not deleted — see §7.
2. Prior art, and what it already settled
Read in full before writing this. Not re-derived.
#354 (shipped, v0.116–0.124) — the £90,000 nudge. A one-way advisory strip driven purely by trailing-12-month income. It does not record whether the business registered, and cannot: it has no input from the user at all.
#829 (shipped, decision-only — outputs/docs/decisions/829-expense-capture-vat.md) — capture a net/VAT/gross split on expense lines: yes. MTD submission: no, permanently. Built as #901. Explicitly recorded as "not blocked on #826 … ship it UK-only now, same as the existing nudge".
#826 (open, Later, effort: high, sub-issue of epic #811epic-selling-outside-uk) — checked as the issue asked. It is not a non-UK-VAT issue; it is the locale/currency layer (country + currency on business_settings, a currency-aware formatter replacing gbp()'s 99 call sites, Stripe EUR price IDs, parse-receipt's £-only assumptions). Unstarted. Its own closing line matters here: "Do not ship a country field before there is something behind it."
#1045 / outputs/docs/invoice-quote-designer-design.md §11 — the issue did not name this, and it is the closest prior art there is. §11's costed alternative (b) is, in as many words, "business_settings.vat_registered + vat_number … and a per-product rate (standard/zero, a CHECK and a drift registration)" — the exact question #1110 asks. Dave answered it on 2 Sep 2026: (a), no VAT line in v1, with (b) "recorded in §11 as the costed alternative". That answer was scoped to v1 of the document, not to the setting in general, and #1110 was filed nine hours later. This document is where (b) gets its full answer; it does not overturn (a).
accounting-capability-research.md — §2 (the £90k threshold and the zero-rated trap), §3 item 8 (the nudge, and #829's correction to it), §4a/§4c (MTD, and why recognition is not worth pursuing), §5a (VAT invoicing is FreeAgent's job), §7 items 9 and 10.
3. The tax facts that decide the shape
Verified 3 Sep 2026 against HMRC guidance and professional-body sources (GOV.UK direct fetches are proxy-blocked from a session sandbox — same sourcing caveat as accounting-capability-research.md §2). Sources at the foot.
Registration threshold £90,000 of taxable turnover in any rolling 12 months; deregistration £88,000. Both unchanged since 1 April 2024 and current for 2026-27. (A reduction has been speculated about in the trade press. It is not in force — watch, don't build. The figure is already data in one place, so a change is the one-line update #354 designed for.)
Zero-rated sales still count toward the threshold. This is the trap #354 exists to warn about and it remains correct.
🔴 "Small food business" is not uniformly zero-rated, and this is where the issue's framing needs sharpening. Under VAT Notice 701/14:
Zero-rated: most cold food for human consumption — plain bread, rolls, cakes, most cold bakery sold to take away.
Standard-rated (20%):all catering, all hot takeaway food, confectionery (sweets, chocolate bars, biscuits wholly or partly covered in chocolate), ice cream, savoury snacks, and anything consumed on the premises.
So a cold-cake maker is zero-rated; a caterer, a hot-food market stall, a deli with seating, or someone selling chocolate-dipped biscuits or brownies-as-confectionery is standard-rated on some or all of their output. ProvenBatch's stated customer is any small food business — caterers, butchers, delis, farm shops, market stalls, jam makers — so the app cannot assume its users' output is zero-rated. The borderline (freshly-baked vs hot takeaway) is genuinely difficult and is not something ProvenBatch should try to adjudicate.
A wholly zero-rated business can register voluntarily and is then in a permanent repayment position — it charges 0% output VAT and reclaims input VAT on ovens, packaging, fuel and standard-rated ingredients. This is a real and common reason a small food business registers well under £90,000. Consequence for the data model: registration status is independent of turnover, so the flag must be user-set and must never be inferred from the #354 nudge.
VAT-registered businesses must keep records 6 years. The app's receipt-retention default is already 6 years and labelled "HMRC-aligned".
Registering makes MTD for VAT mandatory — digital records, filed through software. See §7 for why that is not our problem.
4. Touch-point enumeration — where registration would change behaviour
Every plausible place, with an honest verdict. "No change" is the right answer in several of them and is stated rather than padded.
4.1 The #354 threshold nudge — ✅ REAL CHANGE, and the only live wrongness today
Today the strip fires on trailing income alone. Above £90,000 it reads "you've passed the £90,000 VAT registration threshold. Registration deadlines apply" — which, to a business that registered last year, is a permanent, un-dismissable, factually irrelevant nag on the tab they use for their books. There is no way to make it stop, because the app cannot be told.
With the flag: suppress the "you must register" framing for a registered business (either drop the strip entirely, or replace it with a neutral line). This is the single clearest win and it is one if.
⚠️ One honest caveat on urgency. The strip only appears from £72,000 (80% of £90,000) of trailing 12-month income. If no business on either database is near that, this defect has never actually fired for anyone, and even this stage is pre-emptive. That is checkable from the admin console and should be checked before the work is scheduled — it is the difference between "fix a live annoyance" and "build ahead of demand". It does not change what to build, only when.
4.2 Product pricing, margin and batch economics — ✅ REAL CHANGE, and the biggest one
products.unit_price is captured as "Sale price each £ — what one sells for", and drives productMargin(), the Margin % column, and batch economics (sale_price_each). Nothing anywhere states whether that figure includes VAT.
Unregistered business: no VAT exists. The figure is unambiguous and every number is right.
Registered, zero-rated output: output VAT is 0%, so gross = net. Every number is still right.
Registered, standard-rated output (the caterer, the hot-food stall, the confectioner): the price a customer pays contains 20% that belongs to HMRC. Treating it as revenue overstates margin by up to 20% of the sale price, and overstates Income and Net profit with it.
This is a genuine correctness defect for a real slice of the target market, and it is the strongest argument that VAT-registration is a data-model question and not a cosmetic one. A vat_registered boolean alone does not fix it — fixing it needs a per-product liability and an explicit VAT-inclusive/exclusive convention. It is recorded here so the next person to open this does not have to rediscover it.
4.3 The invoice and the quote — ✅ REAL CHANGE, already decided for v1
A VAT-registered business issuing an invoice is legally required to issue a VAT invoice with specified particulars: its VAT registration number, the rate per line, the VAT amount, and a tax point. orderDocHtml() prints none of that, deliberately (#1045).
The current posture is defensible precisely because it is silent rather than wrong: a document that omits VAT is a document the business must replace, whereas one that invented a rate would be worse. The USER-GUIDE says so plainly (line 1922). Dave's (a) answer stands for v1. This document does not reopen it; it records that the flag is the prerequisite for ever revisiting it, and that revisiting it needs vat_number and per-product liability, not just the boolean.
4.4 Expense VAT capture and the "VAT paid" tile (#901) — ⚠️ WORDING ONLY, no gating
The capture itself is correct for both kinds of business and must stay that way — #829's reasoning was explicit that input-VAT visibility is "useful now, for the business that never registers", because VAT paid on purchases is a real cost either way.
What registration changes is meaning, not data:
Unregistered: VAT paid is an unrecoverable cost, already inside the expense total. Correct as is.
Registered: VAT paid is reclaimable input tax — and under cash-basis bookkeeping it should arguably not sit inside the expense total at all, because it comes back.
🔴 Do not gate the tile on the flag.businessTracksVat() shows it when there is VAT data to show, which is the better trigger: it appears when it is useful and stays away otherwise, and #901 deliberately asked the whole cache rather than the visible month so it could not appear and vanish as you page. Driving visibility from a setting would make a registered business with no recorded VAT see an empty tile, and an unregistered business that records VAT (correctly) see nothing.
The honest change here is one word of framing — call it reclaimable for a registered business — and it is low value. Worth doing only alongside something else.
Update 4 Oct 2026 (#3082): done, once #1147 put income on a net basis and made the mismatch visible. The tile is still gated on businessTracksVat() as above; only its name and the expense totals follow the flag.
4.5 Accounts income, other income and daily takings — ⚠️ INHERITS 4.2, no separate change
Income is recorded gross: order price, plus #346a's other income / daily takings. For an unregistered business, or a registered zero-rated one, gross is the whole of income and this is right. For a registered standard-rated business it includes output VAT that is not theirs. Same defect as §4.2, same fix, not a separate one — listed so the enumeration is complete.
4.6 The accountant export and SA103 mapping — ⚠️ INHERITS 4.2 and 4.4
SA103 turnover for a VAT-registered business is normally reported net of VAT. The export already carries the input-side Net/VAT/rate columns from #901; the output side would need §4.2 resolved first. The flag alone changes nothing here — at most it could add a line to the export's README saying which basis the figures are on, which is worth doing whenever §4.2 lands.
4.7 Receipt retention — ❌ NO CHANGE NEEDED
VAT-registered businesses must keep records 6 years. The default is already 6 years and labelled "HMRC-aligned"; the setting already allows longer. A registered business that has manually set 3 years is the only gap, and a nudge for it is not worth a column. Already correct.
4.8 Labels and PPDS — ❌ NO CHANGE, AT ALL
Verified against buildLabelContent(): a label carries no price of any kind. Name, ingredients, QUID, allergens, dates, storage, net quantity, business name and address — no money. UK PPDS / FIC labelling has no price requirement, and price marking is a separate regime that has nothing to do with VAT status.
VAT-registration status changes nothing about any label ProvenBatch prints. Stated plainly because "VAT should be managed consistently across the app" could otherwise be read as reaching the compliance surface, and it does not.
The shopping list carries supplier prices, but it is an internal purchasing document, not a customer-facing one, and the figures are what the business pays — the input side, already handled by #901. Nothing else in these surfaces touches money.
4.10 Our own plan pricing (PRICES_GBP, upgrade, Stripe) — ❌ OUT OF SCOPE
This is Stella Apps' own VAT position (D-16, amended 31 Jul 2026 — "the price you see is the whole price"), not the customer's. Unrelated to a business_settings flag and explicitly not touched by #829 either. Named here only so nobody conflates the two.
4.11 Non-UK (#826) — ❌ NOT A BLOCKER
#826 is the locale/currency layer, is unstarted, Later, effort: high, and sits under epic #811. Every VAT constant in the app is already UK-shaped by decision (VAT_REG_THRESHOLD, VAT_RATES), on the explicit reasoning that rates are legislation and #826 generalises later if ever needed.
Ship UK-only. Do not hold a UK win for a locale layer nobody has asked to build — the same call #829 made, for the same reason. #826's own warning ("do not ship a country field before there is something behind it") cuts the same way: a country column would be the speculative one here, not vat_registered.
5. The data-model question, answered
The issue asks: is a single business_settings.vat_registered boolean enough, or does #829's VAT-rate capture need a home on the business record too?
Neither, exactly. The boolean is enough for what can honestly be built today, and #829's rate capture should not be promoted to the business record.
Why a default output VAT rate on the business record is the wrong shape. It is tempting — one column, mirrors VAT_RATES — and it would be wrong, because liability is a property of the product, not of the business (§3). The same trader can sell a zero-rated loaf and a standard-rated toasted sandwich on the same day; a business-level default would be right for one and silently wrong for the other, on a legally-loaded figure. #901's rates belong where they are: on the purchase line, describing what a receipt actually printed.
The staged recommendation
Stage 0 — now, if §4.1's caveat checks out. One column.
business_settings.vat_registered boolean null -- null = never asked, distinct from false
Nullable three-state on purpose.null (never answered) must be distinguishable from false (told us they are not registered), or the nudge cannot tell "not registered" from "hasn't said".
⚠️ Must join the column-level grant update (…) allowlist in its own migration, or the write silently fails — the #558/#559/#560 trap, caught three times by bsgrants560.js.
No CHECK constraint, so no drift_info() pair and no check-drift.command registration — shaped like #901's columns, not like an enum.
UI: a single control in Settings → Edit business details, in a new VAT section directly under "Food business registration" (which is the closest existing neighbour — it is the other "are you registered with an authority" fact). Not a bare checkbox: a three-way that can stay unanswered.
Its only behavioural consumer is §4.1's nudge. Plus the USER-GUIDE sentence split in §7.
Deliberately NOT included: vat_number. It has nothing to print until §4.3 is revisited, and #826's own lesson is not to ship a field with nothing behind it. It joins at stage 1.
Stage 1 — awaiting evidence. The output-VAT layer.
Everything that actually needs modelling: business_settings.vat_number; a per-product liability (standard / zero / reduced, with a CHECK, and therefore a drift_info() pair and a drift registration); an explicit VAT-inclusive-or-exclusive convention for unit_price with the migration to match; §4.2's margin and economics maths; §4.5's income aggregation; §4.6's export basis; and §4.3's VAT invoice particulars. A migration on both databases, a UI pass across products, orders, Accounts and the document, and a real risk of getting a legally-loaded number wrong.
The evidence that unlocks it, named as WAYS-OF-WORKING §1 requires:a real business tells us it is VAT-registered and wants to issue invoices or see margins from ProvenBatch. Today the customer base is overwhelmingly under £90,000 and selling zero-rated cold bakery, for whom stage 1 changes nothing (§4.2). Nobody has asked. Building it now would be the largest speculative build in the Accounts area, on the one subject where being wrong has legal consequences for the customer.
Never — see §7.
Why not just do stage 1 now, since stage 0 is small?
Because stage 0 is genuinely independent and genuinely cheap, and stage 1 is neither. Stage 0 delivers the whole of §4.1 with one nullable column and one if, is reversible, and collects the fact we would need anyway. Stage 1 without a customer to check it against would mean choosing the VAT-inclusive/exclusive convention, the liability vocabulary and the invoice layout in the dark — and getting the convention wrong is a silent data migration later, not a UI tweak.
Why not defer stage 0 too?
Reasonable, and the honest answer depends on §4.1's caveat. If a business is already over £72,000, the strip is actively wrong for them and stage 0 is a fix. If none is, stage 0 is cheap preparation rather than a fix, and can wait for the same evidence as stage 1. Check first.
6. What this means for the #354 nudge specifically
Recorded separately because the issue asks directly whether the nudge should change wording or stop.
Registered (true): the "you must register" framing is wrong. Suppress it. A registered business does not need a threshold warning at all — they are past the decision. (Not "keep it but soften it": there is nothing left to warn about.)
Not registered (false): the strip is exactly right and must keep firing, including the zero-rated-sales trap sentence, which is the whole point of #354.
Unanswered (null): keep today's behaviour unchanged. The strip is the prompt that would make someone answer, so it must not require the answer it is trying to elicit.
Do not infer the flag from the nudge. §3's voluntary-registration point means a business well under £72,000 may be registered, and one over £90,000 may not yet be. They are independent facts and the app should treat them that way.
Keep VAT_REG_THRESHOLD and VAT_NUDGE_AT as data in one place — #354's design, and the reason a threshold change is a one-line update.
7. 🔴 MTD submission stays out — restated, not reopened
outputs/docs/decisions/829-expense-capture-vat.md and accounting-capability-research.md §4a/§4c settled this. Restated here because a vat_registered flag is precisely the change that invites someone to reopen it — the moment the app knows a business is registered, "so shouldn't we file their return?" becomes the obvious next thought. It is answered:
Registration makes MTD for VAT mandatory for the business, not for us. The obligation is the customer's, and it is discharged through their accountant or bridging software.
A CSV export is an explicitly valid digital link (VAT Notice 700/22 s4.2.1). The only banned link is copy-and-paste re-keying. So a clean export is a legitimate part of functional compatible software, with no HMRC integration of our own.
HMRC recognition is not worth pursuing (§4c): fraud-prevention headers on every API call as a permanent legal obligation, an ongoing validator, a change-log — a compliance workstream orthogonal to what makes ProvenBatch valuable.
The truthful phrasing never changes:"keeps the digital records; exports them for your MTD software or accountant" — never "MTD compatible" or "MTD recognised".
Nothing in this document reopens that, and nothing built from it should. If stage 0 ships, the USER-GUIDE sentence at line 2037 must be split, not deleted: the app may then know whether you are registered, and it still does not compute or submit a VAT return, and still isn't going to.
8. Recommendation
Build stage 0 — business_settings.vat_registered, nullable three-state, on the column-grant allowlist, surfaced in Settings → Edit business details under a new VAT heading, consumed only by the #354 nudge (§6) and the USER-GUIDE split (§7). First check whether any business is over £72,000 trailing income (§4.1); that decides whether this is a fix or preparation.
Do not promote #829's VAT rate to the business record. Liability is per-product, not per-business (§5).
Hold the output-VAT layer as awaiting evidence, with the gate named: a real business tells us it is registered and wants invoices or margins from ProvenBatch (§5).
Record §4.2 as a known defect — margins and income overstated for a registered, standard-rated business — so it is not rediscovered. It is the strongest single argument for stage 1 when the evidence arrives, and it belongs in the stage-1 issue's Description.
Ship UK-only. #826 is not a blocker (§4.11).
Do not touch labels (§4.8) or D-16 (§4.10).
Do not reopen MTD (§7).
Follow-up issues this record proposes (none raised by it — that is Dave's call, per #1045's pattern):
(Accessed 3 Sep 2026. GOV.UK direct fetches are proxy-blocked from a session sandbox, so GOV.UK content arrived via search excerpts cross-checked against accountancy-practice summaries — the same caveat as accounting-capability-research.md §2. Anything below deserves a direct GOV.UK read before being hard-coded into a build.)
MTD (unchanged, quoted from the existing record rather than re-researched): VAT Notice 700/22 s4.2.1 (digital links); accounting-capability-research.md §4 and its source list.
Evidence
Last updated 17 Jul 2026
5. Evidence from our own data
The most valuable research here isn't the market — it's what our live data revealed about the product's assumptions.
Finding
Date
Why it matters
Dark Chocolate Chips flagged "Milk: contains" — and human-verified — with a declaration of "Cocoa Mass, Sugar, Cocoa Butter, Fat Reduced Cocoa Powder, Emulsifier (Soya Lecithins)". No dairy. \bbutter\b matched "Cocoa Butter".
17 Jul 2026
Human verification is not a backstop. Someone ticked this. Assume every manual check has a false-negative rate. Fixed v0.11.0.
Kinder Bueno had zero allergens flagged despite milk, wheat, hazelnuts and soya in its text. Root cause: allergens were only ever detected during the original workbook import — nothing scanned anything added in the app since.
17 Jul 2026
A silent, dangerous under-declaration produced by an invisible architectural gap. Fixed v0.12.1 (#44).
38 of 63 allergen links unverified, across 24 ingredients.
17 Jul 2026
The honest state of auto-detection. Now visible in Allergens.
Only 6 of 73 products have recipe quantities.
17 Jul 2026
The allergen matrix can only cover 6 products. Data entry — not features — is the bottleneck. Reinforces #27.
Generic "nuts" was never detected — old and new term lists knew only specific nuts (almonds, hazelnuts…). Packs say "may contain nuts" constantly.
17 Jul 2026
Fixed v0.11.0. Shows keyword lists fail on the common case, not the exotic one.
"Nougat → nuts" would have been a false positive — industrial nougat is sugar/glucose/palm fat/barley malt/milk.
17 Jul 2026
Richer word lists manufacture new errors. Any check must flag for review, never auto-correct.
OFF's ingredients_text is in the product's own language, and ingredients_text_en is sometimes not English either. A French list scans as containing no allergens.
17 Jul 2026
Importing foreign text is worse than importing nothing. Fixed v0.12.2 (#8).
The pattern: every serious defect found so far has been silent under-declaration caused by something invisible — not a crash, not a wrong number. Design accordingly: prefer loud, additive, reviewable over clever and quiet.
Evidence
Last updated 24 Aug 2026
17. Onboarding friction: why one-at-a-time pack reading does not clear a backlog (researched 24 Aug 2026)
Raised by Dave, 24 Aug 2026: could a user photograph a batch of declarations and have the app read them all in at once? Design: outputs/docs/batch-pack-intake-design.md. Mock-ups: outputs/docs/batch-pack-intake-mockups.html. Backlog: #673 (parent), #674, #675, #676.
17a. The finding that matters most: we already fixed this, and it did not work
§2 records the adoption barrier as evidenced rather than assumed — ~70 declarations typed by hand. #249 (v0.73.0, late Jul 2026) was the answer: photograph one pack, parse-pack transcribes the declaration verbatim with Opus vision, compare side by side, confirm. It works. Measured on production 24 Aug 2026, roughly four weeks after it shipped:
Signal
Value
supplier_product_sourcings
181
still carrying the (enter declaration) placeholder
91
never confirmed against a pack (declaration_checked_at is null)
157
pack_photos rows since #249
19
api_usage_log rows for parse-pack
16
pack_parse_events rows (#553 Phase A)
1
brand_products rows (#553 Phase B library)
0
measured cost per pack read
£0.0161 (mean of 16)
Half our own catalogue still has no declaration at all, on the founder's instance, with a working one-tap reader shipped. This supersedes §2's "~70 sourcings" figure — not because §2 was wrong, but because it is now measurable after the intervention rather than before it. §2's conclusion stands; its arithmetic is stale.
Two second-order findings fall out of the same query:
brand_products is empty, so #553's prefill banner and its parse-pack brand hint both currently do nothing. The library needs a volume source and per-item capture is not one.
declaration_checked_at went from 0/130 to 24/181 in the month. Real movement, and still an eighth of the catalogue. The per-item flow is not the bottleneck being solved; it is the bottleneck.
17b. The structural blocker, in one line
pack_photos.sourcing_id is NOT NULL, and capturePack() refuses outright with "Add a supplier product first; a pack photo has to belong to one". You must build the record before you may photograph it. For a brand-new business with an empty app that is exactly backwards, and it is why the shipped feature cannot serve the onboarding case it was aimed at.
17c. Claude vision limits, checked against the current docs (24 Aug 2026)
Checked rather than recalled, because the batching question turns on them:
Up to 600 images per API request on a 1M-context model (100 on a 200k-context model such as Haiku 4.5); 20 per turn on claude.ai.
⚠️ **Above 20 image blocks in one request, a stricter per-image dimension cap applies to every image in that request.** Stay at ≤20 blocks, or resize so neither dimension exceeds 2000 px.
10 MB per image (base64) on the first-party API; 32 MB per request, which is usually reached before the image count is.
Cost is ⌈width / 28⌉ × ⌈height / 28⌉visual tokens. Claude 4.7 and later are a high-resolution tier: max long edge 2576 px, capped at 4784 visual tokens; earlier models 1568 px / 1568 tokens. Oversized images are downscaled, not rejected.
The Batch API is 50% cheaper and asynchronous, up to 24 h.
Consequences for us.compressReceiptImage's existing 1600 px / q0.8 output already sits inside every one of these limits, including the >20-block dimension rule, so no new sizing is needed. And batching many packs into one request buys nothing: tokens are tokens, so the price is identical, while a single request spanning forty different packs invites attaching a declaration to the wrong pack. §5 records that every serious defect found in this product so far has been silent under-declaration. One pack per call preserves the per-pack legible gate and confidence, which is the whole safety mechanism. The Batch API is the wrong trade for an interactive onboarding: 24 h latency and a poller, to save pennies.
17d. Our own rate limits are the binding constraint, not the model
DAILY_CALL_CAP = 20 per user per slug per day. A 70-pack sweep takes four days.
DAILY_SPEND_CEILING_GBP = 10, and spendTodayGbp() filters on slug and date only, with no user filter — so it is a global ceiling across every tenant, not a per-tenant one. At £0.0161 a read, ten beta testers onboarding 70 packs each on one day is ~£11.30, and they would lock each other out. Worth deciding before beta rather than discovering in December.
17e. Where the market is on this
Barcode import is the competitors' answer, and it is the one we already rejected. FoodCore ships Open Food Facts barcode import (#432). Their own page carries the caveat verbatim: "it's worth reviewing imported allergen information against the physical packet, particularly for products with 'may contain' warnings which may not always be captured in database records." Their admission is evidence our v0.75.0 removal was reasonable, not evidence it is now safe to reverse. Supermarket own-brand is about half a small food business's shopping and that data is the retailer's own asset.
The serious UK data sources are wholesale infrastructure. Erudus covers 200+ attributes across ~25,000 products with a real API, and is sold to manufacturers, wholesalers and foodservice caterers, not to a market-stall jam maker. Nielsen Brandbank and GS1 productDNA are the same shape. The gap is structural, not a data-quality wrinkle that improves — which is exactly what #249 concluded, and it still holds.
The competitive framing is unchanged and in our favour. Everyone in this space names ingredient-library setup as the reason businesses delay switching. Their answer is a barcode lookup that misses own-brand. Ours reads the pack in the user's hand, which always exists.
17f. The rule that shapes the product, and it is already written down
DESIGN-LANGUAGE.md § Bulk select & edit (#437): "Compliance data is never bulk-set — a bulk 'set declaration' or 'set allergen' would be a false-declaration machine."
So a batch flow can batch capture and reading, and must not batch acceptance. That is not a limitation to design around; it is the correct answer, and the value survives it intact. The time in the current flow goes on navigating to forty records and waiting for forty readings, not on the forty judgements themselves.
17g. One thing we have never said out loud
Receipts and packs are complementary. A receipt photo names ~30 supplier products in one shot and carries no declaration. A pack photo carries the declaration and names one product. The strongest onboarding is scan the receipt to build the skeleton, then sweep the shelf to fill the declarations. Both readers have been live since M1. All that is missing is saying so at the moment a new business has an empty app.
Sources:
Live production SQL, 24 Aug 2026 (supplier_product_sourcings, pack_photos, pack_parse_events, brand_products, api_usage_log) — the table in §17a is measured, not modelled
supabase/functions/parse-pack/index.ts — DAILY_CALL_CAP, DAILY_SPEND_CEILING_GBP, spendTodayGbp(), the 4-image slice and the 8 MB guard, read from source
27. GitHub Actions cost — where 38,000 billed minutes a month go (researched 13 Sep 2026)
The finding: three quarters of the Actions bill is ship.yml, and most of that is two recent changes whose billing effect nobody measured. Whole history pulled through the API — 14,569 runs in the 30 days to 13 Sep ≈ 38,000 billed minutes against 3,000 included (~$210 overage), with the run-rate tripled since 3 Sep. Full write-up, method and the implementation order: outputs/docs/actions-cost-research-2026-09.md.
Workflow
Runs
Est. billed min
Share
ship.yml
2,177
~29,050
76%
land.yml
661
~5,200
14%
project-status.yml
2,620
~1,180
3%
land-sweep.yml
748
~770
2%
everything else (52 files)
~1,700
~1,800
5%
Two causes. (1) GitHub bills every job rounded up to a whole minute, so the six-way shard matrix (#1154) bills six two-minute shards as 18 minutes where one six-minute job billed 7 — a PR push went from ~11 to ~24 billed minutes. (2) The Gated-by: land.yml receipt that lets a landing skip the re-gate on main misses on the #962 fast-forward path (a docs push landing mid-gate puts a receipt-less merge commit on HEAD): 42 of the 51 landing pushes since 3 Sep ran the full suite again. Behind those: four always-on jobs billing 4 minutes for ~30 s of work on every event (~8,700 min/month), and ~4,000 min/month of one-step jobs paying the one-minute floor that the ubuntu-slim runner ($0.002/min vs $0.006) exists for.
Fewer version numbers is NOT a cost lever. 313 versions were stamped in the window against ~98 tag bursts; the promote is already the batched act and a stamp costs seconds.
Decided (Dave, 13 Sep): re-shard 6 → 3 and draft PRs skip the suite; the release model moves to version-at-promote / build-at-landing as its own epic. Filed 13 Sep as #1784 epic-actions-cost (eight children) and #1785 epic-release-at-promote (five, sequential). C1 (#1794–#1797) is the code+docs; #1798 waits for one live promote. Estimated effect of the pipeline changes: ~38,000 → ~13,000–15,000 min/month.
Second pass (1 Oct 2026): §35. The bill rose anyway: the suite slowed 3.5× on the 8-core runners and most landings still carry one PR. See §35 and outputs/docs/actions-cost-research-2026-10.md.
Sources: GitHub Docs Actions runner pricing and GitHub Actions billing; GitHub Changelog 16 Dec 2025 (pricing) and 22 Jan 2026 (ubuntu-slim GA); the repo's own run list via the API.
Full write-up
From outputs/docs/actions-cost-research-2026-09.md
GitHub Actions cost — where the minutes go, and the three decisions (review, 13 Sep 2026)
Asked by Dave after another month of buying additional Actions minutes. Scope: measure where the billed minutes actually go, say whether "getting things into staging" can be made cheaper, say whether batching PRs into fewer version jumps before a promote would save money, research what the industry does, and propose fewer, larger releases in place of the point-release stream. Companion to .github/DEPLOYMENT.md (the pipeline as designed) and outputs/docs/multi-session-shipping-review-2026-08.md (the last time the shipping process was measured rather than reasoned about).
Status: C1 is built (#1794–#1797 on this PR); #1798 is blocked until one live promote under the new model. Dave read the findings and made three decisions in the same session (§ What was decided). On his instruction the work was filed the same day as two epics: #1784 epic-actions-cost (Part A + B1: #1786 receipt → #1787 re-shard → #1788 drafts skip → #1789 advisory jobs → #1790 cleanups → #1791 ubuntu-slim → #1792 budget → #1793 re-measure) and #1785 epic-release-at-promote (C1: #1794 → #1795 → #1796 → #1797 → #1798, sequential). C2 and C3 were declined and stay declined. The numbers below are the baseline those epics are measured against.
The short answer
Three quarters of the bill is one workflow, ship.yml, and most of that is two changes made for good reasons in the last two weeks that turned out to cost money in a way nobody measured:
Sharding the suite six ways (3 Sep, #1154) halved wall-clock and doubled the bill. GitHub bills every job rounded UP to the next whole minute. Six shards of ~2 minutes each bill as 6 × 3 = 18 minutes, plus six checkouts and six npm ci; the one job they replaced billed 6–7. A PR push went from ~11 billed minutes to ~24. There are ~1,060 PR runs a month.
The "already gated by land.yml" receipt is not matching most landings, so ship.yml re-runs all six shards on main after the Land gate has just proved the same tree. Measured on the run list: 42 of the 51 landing pushes since 3 Sep paid the full gate again. The cause is the #962 fast-forward path, which puts a receipt-less merge commit on top of the stamped one whenever a docs push lands mid-gate — which, at this repo's HANDOVER cadence, is most of the time.
Behind those two: four always-on jobs that bill four minutes for thirty seconds of work on every single event (~8,700 min/month), and a long tail of one-step event and cron jobs paying the one-minute floor (~4,000 min/month) that could run on a runner a third of the price.
Batching PRs into fewer version numbers would not reduce the bill. The promote is already the batched act — 313 app versions were stamped in 30 days against ~98 tag bursts, with recent hops of 8–23 versions — and stamping a version costs seconds. The gate already runs once per Land batch (#1140). Fewer releases is a process win, and worth doing for its own reasons (§ C); it is not where the money is.
Getting things into staging is already cheap — the staging job itself is ~40 seconds. What is expensive is everything that runs around it on the same push: the re-gate that should have been skipped and the four always-on jobs. Fix those and a landing reaches staging in ~2 billed minutes instead of ~21, about four minutes sooner.
Measured baseline: ~38,000 billed minutes in the 30 days to 13 Sep against 3,000 included, so roughly $210 of overage at $0.006/min — and the run-rate tripled after 3 Sep (ship.yml went from ~490 to ~1,830 billed min/day), so September is heading for $400–600 if nothing changes. Part A below takes that to an estimated 13,000–15,000 min/month (under $75 of overage) with no loss of coverage; the runner move in A4 shaves the rest.
The evidence
Method
The whole Actions history was pulled through the API (list_workflow_runs: 15,864 runs since 22 Jul 2026; 14,569 in the window 14 Aug → 13 Sep). get_workflow_run_usage now returns billable.UBUNTU.total_ms: 0 for every run — GitHub retired per-run billable data — so billable minutes were calibrated from per-job timestamps (list_workflow_jobs) using GitHub's own rule: each job rounds up to the next whole minute, skipped jobs are free ([runner pricing][pr]). The model was checked against nine fully enumerated ship.yml runs (within ±3 minutes of the exact per-job sum on each). Every runner in the repo is ubuntu-latest except the dispatch-only ios-preview.yml (macos-latest, billed 10×, 0 runs in the window). Numbers below are estimates from that model, good to about ±10%.
Top five = 96%. Of the 14,569 runs, 8,667 (59%) were skipped runs that bill nothing — almost all feedback-status-sync.yml waking on every issue comment and exiting. They cost money only in run-list noise.
Cancelled runs cost ~1,380 min (103 cancelled sharded PR runs at ~16 min each — all six shards had already started when the newer push arrived — plus 54 cancelled Land runs). Failed ship.yml runs cost another ~1,380. Re-runs and queue waits are negligible (13 runs with run_attempt > 1; job created → started is 1–3 s throughout).
What one event costs today, job by job
ship.yml, PR push, current shape (run 34690906323, wall 3m54s, 15 jobs): classify 19 s, journeys 22 s, versions 16 s, ruleset 4 s, suite (1..6) 1m51s–2m27s each, test 24 s, preview 33 s → 23–24 billed minutes. The same PR push before 3 Sep (run 32958288792, wall 6m49s): classify 10 s, versions 7 s, ruleset 3 s, test 6m09s, preview 20 s → 11.
ship.yml, push to main after a landing, current shape (run 34493765784 "Merge PR #1616…", wall 4m12s): the four small jobs, suite (1..6) 1m39s–2m35s, test 26 s, staging 38 s → ~21 billed. The same push when the receipt does match (run 34387935516, wall 1m44s): four small jobs + staging 1m18s → 6. A documents-only push (run 34000153149, wall 19 s): four jobs of 12–16 s → 4 billed minutes.
land.yml (one job, unsharded full gate): quartiles 7.9 / 11.4 / 13.1 billed minutes. promote.yml: 1–2 billed minutes — checkout, drift gate, probe, deploy, tag, release; the suite is inherited, not re-run.
Per PR
500 distinct PR branches had ship.yml runs; 1,061 runs → 2.12 runs per PR (median 1; 124 PRs had three or more runs and account for 570 of them; five branches had 14–15 runs each).
The price list, so the arithmetic is checkable
Rate
Source
Linux 2-core (ubuntu-latest)
$0.006/min (was $0.008 until 1 Jan 2026)
[runner pricing][pr], [changelog][cl-price]
Linux 1-core (ubuntu-slim)
$0.002/min; 1 vCPU, 5 GB RAM, container not VM, 15-minute job cap; GA 22 Jan 2026
[runner pricing][pr], [changelog][cl-slim]
Included minutes, Pro and Team
3,000/month (Free: 2,000); cannot be spent on larger runners
[billing][bill]
Rounding
"GitHub rounds the minutes and partial minutes each job uses up to the nearest whole minute"
[runner pricing][pr]
The repo is private on a personal account, at least Pro (the main protection ruleset needs it — .github/DEPLOYMENT.md § Require status checks), and the org move planned in outputs/docs/github-org-migration-plan.md puts it on Team: same 3,000 minutes. Nothing in the repo records the plan, an Actions budget, or a spending limit.
Five findings
1. Six shards bill twice what one job did — ship.yml:484-519
The 3 Sep measurement that justified the matrix was right about wall-clock (12m35s → ~4 min) and silent about billing. Per-job rounding makes shard count a cost dial, not a free speed dial. Shard compute today is ~9.5 min in total (six shards at 1m50–2m30 including ~35 s of checkout and npm ci each; the suite has grown from 314 files when the shards were sized to 461 now):
Shards
Each shard, approx.
Billed for the suite
Wall
6 (today)
2m00–2m30
18
~4 min
4
~3m00
16
~4 min
3
~3m50
12
~4.5 min
2
~5m45
12
~6.5 min
1
~11 min
12
~11 min
Three shards keep almost all of #1154's wall-clock win at two thirds of its price; two shards are safer against a slow runner tipping a shard over a minute boundary. The guard outputs/tests/shardgate1154.js:62 hardcodes const N = 6 and should read the matrix from the parsed YAML (its own #969 lesson) so it follows the file rather than pinning it.
2. The Land receipt is not on HEAD most of the time — land.yml:505-518, :958-976; ship.yml:132-190
ship.yml skips the suite on a push to main only when HEAD itself carries Gated-by: land.yml run N and the committer is the landing robot — correct, and deliberately never an ancestor. land.yml writes the receipt on every stamp commit. But two of its own commit shapes end a landing without one:
the #962 fast-forward merge (Merge remote-tracking branch 'origin/main' into land-attempt), created when a documents-only push lands during the gate — land.yml:516-517 says "deliberately NOT re-applied", so that landing "pays the full gate on main";
a changeset-less PR (or a batch whose last PR has no changeset), which ends on Merge PR #N (…) into the landing tree.
main as this review is written is exactly shape one: 49f4223b (ff merge, no receipt) on top of 019a0d2a (v0.369.1, receipt present). With HANDOVER pushed at the end of every session and prose going straight to main by rule (§3.1), a docs push landing inside a 10-minute gate window is the common case, not the edge case the comment assumed. Cost: ~21 billed minutes per affected landing and staging ~4 minutes behind it, on ~80% of landings.
The fix is small and does not touch the ancestor rule: after the gate is green and before the push, if HEAD lacks the receipt, git commit --amend its message to add it. HEAD at that point is always a commit this run created and has not pushed. The tree it licenses is the gated tree plus documents-only movement that land-race.sh has already classified through docs-only.sh — the identical claim the docs fast path makes with no gate at all, so nothing is weakened. The guards that drive real git history for these paths (outputs/tests/landdocs962.js, outputs/tests/landbatch1140.js) should assert the pushed HEAD carries the receipt.
3. Four always-on jobs are almost pure rounding tax — ship.yml:57, 310, 365, 435
classify (19 s), versions (16 s), journeys (22 s), ruleset (4 s) run unconditionally on every event: 4 billed minutes × 2,177 runs ≈ 8,700 min/month, about a quarter of the whole bill, for work totalling about a minute. Two of them (classify, journeys) do fetch-depth: 0 clones of a 311 MB .git. On the ~900 documents-only pushes a month they are the entire run.
versions and ruleset are advisory (green-with-warning, never red) and can become steps inside classify with continue-on-error: true. journeys cannot: it goes red on a Tier-1 journey rise, and a red classify would skip suite → test → report SUCCESS to the ruleset (the #22 trap). It stays a job, gated if: github.event_name == 'pull_request' — its comment step is PR-only already and a red on main gates nothing. Result: a PR event bills 2 for these, a push bills 1. Guards to update in the same commit: outputs/tests/versions628.js:105-130 (asserts a versionsjob with no needs — the intent, "can never hold a deploy", survives as a step) and outputs/tests/journeys873.js:167-176.
4. The one-minute floor on one-step jobs — and a runner a third of the price
project-status.yml (1,166 billed runs × ~15 s), land-sweep.yml (734 × ~10 s), feedback-status-sync.yml, epic-children-check.yml, delete-branch.yml, release.yml, and the ~24 daily and weekly crons: roughly 4,000 min/month, nearly all rounding. Every one is a single actions/github-script or node step with no npm ci, no wrangler, no playwright, no apt — which is exactly the shape GitHub's ubuntu-slim runner exists for (node, git, gh, jq, curl preinstalled; 15-minute cap). Moving them is a one-line runs-on: change each, at a third of the rate. outputs/tests/landsweep944.js:355 pins runs-on: ubuntu-latest for the sweep and needs updating. Not movable: anything that runs npx wrangler, npx supabase, playwright, ffmpeg, or the suite.
5. Fewer version numbers is not a cost lever
Facts, from the Releases API and outputs/VERSIONING.md: 313 app versions and 48 admin versions stamped in 30 days (~10 customer-facing release-note paragraphs a day); ~98 tag/release bursts in the same window; hop sizes have grown from 1–4 to 8–23 (the 8 Sep promote carried 23 versions; a 12 Sep one another 23). Production is v0.360.5 while main is v0.369.1.
What that costs in Actions: the stamp is seconds inside a Land run that gates once per batch (#1140); promote.yml bills 1–2 minutes; each tag fires a ~1-minute release.yml run (69 in the window). Collapsing 313 versions into 15 would save a few hundred minutes a month. The reason to do it is the process — the paperwork, the 23-line promote briefs, the #787/#790 tag walk that exists only because there are intermediate versions to tag — and that is § C.
Also found on the way (not cost, but wrong and worth fixing)
tag.yml:39-42 and outputs/docs/handbook/shipping.md (§ what would notice if a workflow stopped firing) both say a tag created inside Actions cannot fire release.yml. That was true with GITHUB_TOKEN; since #1458 both tag.yml and promote.yml push tags with GH_MERGE_TOKEN, and the run list shows 69 release.yml push-event runs in the window — one per tag. Harmless (it finds the Release already made and edits it), but the prose is stale and the handbook's "measured: 10 push runs ever" predates the change.
ship.yml:484-490 and .github/DEPLOYMENT.md:35-41 still size the shards off "314 suites, 12m35s"; the suite is 461 files.
deploy-functions.yml:311,313 runs npx --yes supabase@latest inside the per-function × per-environment loop — one CLI download per function per environment, unpinned.
browser-uat.yml:184, deploy-media.yml:131, w1-clips.yml:53 run playwright install --with-deps chromium uncached on every run (and browser-uat then npm i the same package again). actions/cache on ~/.cache/ms-playwright is the standard fix.
setup-node has cache: npm only inside the gate action and land.yml; the wrangler jobs in ship.yml, promote.yml, deploy-site.yml do not.
What is already right, and should not be touched in the name of cost
cancel-in-progress on PR runs (never on main). Job timeouts everywhere in ship.yml. The docs-only fast path, with docs-only.sh as the only classifier — do not add a trigger-level paths-ignore, that is a second allowlist (§3.1, #1407). promote.yml inheriting the gate verdict rather than re-running it. One Land gate per batch. test keeping its exact name and running on a red shard rather than being skipped. The receipt never read from an ancestor. land.yml the only writer to main. The hourly land-sweep cron — the file already made the cost call (land-sweep.yml:37-41) and 81 runs a month is what it costs. Nothing on macOS.
Recommendations
A. The pipeline changes — one PR, in this order, each measurable on its own Ship runs
#
Change
Files
Guards to update
Est. saving / month
A2
Put the receipt on whatever HEAD the gate passed (amend before push)
land.yml:505-518, :958-976; comment at ship.yml:132-160
landdocs962.js, landbatch1140.js
2,000–2,500 min, and staging ~4 min sooner on most landings
A1
Re-shard 6 → 3
ship.yml:509-519; re-measure the comment block; optionally actions/cache on outputs/tests/node_modules keyed on the lockfile
shardgate1154.js (read N from parsed YAML)
~8,500 min
A3
versions + ruleset become steps of classify (continue-on-error); journeys PR-only
ship.yml:57-482
versions628.js, journeys873.js; grep outputs/tests/*.js for versions: / ruleset: first
~5,400 min
A5
Fixed-cost cleanups: pin and hoist the supabase CLI; cache playwright; cache: npm on the wrangler jobs; correct the release.yml prose
runs-on: ubuntu-slim for every one-step github-script / node job — last, one file at a time, each dispatched once and its run read
the ~25 files named in finding 4; classify/journeys only if their full-history checkout stays under a minute on 1 vCPU
landsweep944.js:355
~4,000–5,000 min move to a third of the price
Order matters: A2 first because it is the cheapest and stops the most visible waste; A1 and A3 next because they are the bulk; A4 last because a missing tool on the slim image fails at runtime, not at parse time, and each move wants its own verified run. Together A1–A3: ~38,000 → ~13,000–15,000 min/month at September's volume. No changeset (release machinery, no customer-visible change). Landed via queued like any other PR; .github/workflows/** is never the docs fast path.
Verification for the PR: read list_workflow_jobs on its own PR push and confirm the arithmetic (target ≤ 16 billed for a code PR, 2 for a docs push, and a landing push whose classify says "Already gated" while staging still deploys and verifies). Make each advisory step fail once on a branch (a deliberately stale LIVE-VERSIONS block) to prove it still can — #22.
A6. Billing hygiene — Dave's dashboard, not the repo
Set an Actions budget with alerts under Billing → Budgets and alerts so overage is a notification rather than a surprise, and record the plan and the budget figure in outputs/docs/handbook/shipping.md. Add "set the Actions budget when the org is created" to outputs/docs/github-org-migration-plan.md § Decisions only you can make. Larger runners are not a lever: included minutes cannot be spent on them and they cost more per minute than the work saves. (Dashboard steps are from GitHub's docs as of today, not memory — verify the menu before clicking, per AGENTS.md #19b.)
B. The PR-side gate — decided: re-shard AND drafts skip the suite
The suite runs three times per change: the pre-push hook locally (fast feedback), ship.yml on the PR, and land.yml on the merged tree — the only run that decides anything. Options considered:
B1 — draft PRs skip the suite. Chosen.suite, test and preview gain github.event.pull_request.draft == false. Safe because land.yml's precheck already refuses a draft (WAYS-OF-WORKING §3.2), so a skipped test on a draft can let nothing land. Workers open as draft and mark ready when their own UAT (§2.2) is done; the review-at-PR-open (#1156) reads the diff and does not need the run. 124 of 500 PRs had three or more Ship runs; this is where they go. Process edits: .claude/commands/batch.md worker prompt, .grok/skills/batch/SKILL.md, WAYS-OF-WORKING §3.2a step 1. Estimated ~3,000–4,000 min/month on top of A1.
B2 — drop the PR-side suite entirely; land.yml is the gate. Declined for now. Biggest lever (10,000+ min/month after A1) but a red PR is found only at landing, and a red batch lands nothing and recovers one PR per pass (§3.2a). Revisit after A1–A4 have been measured for a month.
C. Fewer, larger releases — decided: version at promote, build at landing
Today one changeset per PR becomes one version number at landing (land.yml:588-690, outputs/apply-changeset.sh, which hard-refuses two changesets per surface — #585's "one version = one release-note paragraph"). Promote then tags and releases every intermediate version.
Industry practice for a repo shaped like this is the [Changesets][cs] / [release-please][rp] model: changesets accumulate on the main branch; a version is computed and stamped when a release is cut, from all of them at once. The bump level of the release is the highest bump any accumulated changeset asked for; the release notes are the concatenation. Applied here:
C1 — version at promote, build at landing. Chosen. Landing stops stamping APP_VERSION: it merges, gates, pushes, and the PR's outputs/changes/*.md stay on main. So the staging verify step and check-staging-freshness.js cannot pass vacuously (#22), landing stamps a cheap APP_BUILD (short SHA or a counter) that carries no paperwork. Promote becomes "cut a release": consume every pending changeset, compute one version (feature bump if any changeset is a feature, else minor), write one VERSIONING entry and one RELEASE-NOTES section with a bullet per changeset, stamp the USER-GUIDE marker, commit, deploy, verify, one tag, one GitHub Release. The promote brief is the release notes. #787/#790's tag walk goes away. Touches: land.yml, promote.yml, outputs/apply-changeset.sh, outputs/changes/README.md, outputs/check-versions.command and the LIVE-VERSIONS block (staging shows a build id), .github/scripts/check-staging-freshness.js, guards changeset846.js, relnotes585.js, sitecurrency.js, promotetags790.js, versions628.js, WAYS-OF-WORKING §2.3 / §3.2 / §3.4, AGENTS.md #6 / #7, handbook/versioning.md. Non-negotiables preserved: nobody hand-bumps a version (#6 — the stamp just moves from land.yml to promote.yml); LIVE-VERSIONS stays machine-checked (#7); a promote still ships the whole of main and still needs Dave's word on the brief (#10); the tag is still made inside Actions (#11).
C2 — coalesce per Land batch only. Declined. Relaxing apply-changeset.sh so one Land run stamps one version per surface would only cut versions to ~150–190/month, because most landings are singletons.
C3 — keep the numbering, make promote produce one tag and one Release. Declined. Trivial, but the point releases stay.
Suggested issue split for C1, ready for the create-issues skill when Dave asks — an epic plus five children, sequential (each touches the stamping path):
Build id at landing.land.yml stops calling apply-changeset.sh; stamps APP_BUILD; ship.yml's staging verify and check-staging-freshness.js compare the build id; LIVE-VERSIONS gains a build column for staging. Guards: versions628.js, the freshness suite.
apply-changeset.sh consumes N changesets into one version. Drop the one-per-surface refusal; bump = max over the set; one VERSIONING entry and one RELEASE-NOTES section with N bullets; USER-GUIDE marker once. Guards: changeset846.js, relnotes585.js ("one version = one section"), sitecurrency.js.
promote.yml cuts the release. Runs the stamper on main, commits with the receipt so ship.yml does not re-gate the version-string change, deploys the stamped tree, verifies, one tag, one Release; deletes promotetags790.js's walk. Brief = the generated notes.
Documents. WAYS-OF-WORKING §2.3 / §3.2 / §3.4, AGENTS.md #6 / #7, handbook/versioning.md, outputs/changes/README.md, .github/DEPLOYMENT.md, this review's status line.
Retire release.yml's push: tags trigger and tag.yml's multi-version repair mode once one promote has produced one tag end to end.
What was decided (13 Sep 2026, Dave, in the session that wrote this)
PR gate: re-shard 6 → 3 AND draft PRs skip the suite (A1 + B1). Land still re-gates.
Release model: C1 — version at promote, build at landing. Its own epic.
This session writes the review only. No workflow edits, no issues filed. Dave decides what to file after reading it.
What to watch, to know whether it worked
Re-run the measurement a month after Part A lands — the method is in § The evidence, and it is mechanical enough for a session: page list_workflow_runs for the window, list_workflow_jobs on a sample, round each job up. The numbers that should move: a code PR push ≤ 16 billed minutes; a documents-only push 1; a landing push 2 with classify reporting "Already gated"; a site-only PR push ≈ 4 billed and a site-only push to main ≈ 3 (#1957, 18 Sep 2026 — classify plus one unsharded site-suite plus test, with no shards, no preview and no staging redeploy; it was ~12); ship.yml under ~600 billed min/day at September's PR volume; total under ~15,000/month. If a shard sits within 15 s of a minute boundary on three consecutive runs, drop to two shards rather than tune around it. And watch the budget alert — it is the only signal that does not depend on someone remembering to measure.
Sources (fetched 13 Sep 2026)
The headline version of this review is outputs/RESEARCH.md §27; this document is the full write-up the Research library attaches to it.
[GitHub Docs — Actions runner pricing][pr]: the per-minute table, the per-job round-up, included minutes unusable on larger runners.
[GitHub Docs — GitHub Actions billing][bill]: 2,000 / 3,000 included minutes by plan.
[GitHub Changelog, 16 Dec 2025 — pricing update][cl-price]: hosted runners cheaper by up to 39% from 1 Jan 2026; the self-hosted platform charge postponed.
[GitHub Changelog, 22 Jan 2026 — 1 vCPU Linux runner GA][cl-slim]: ubuntu-slim, 1 vCPU / 5 GB, container-based, 15-minute cap.
[Changesets][cs] and [release-please][rp]: the accumulate-then-cut release model in § C.
The repo's own record: land-sweep.yml:37-41 (the only prior cost arithmetic), ship.yml:132-160 and .github/DEPLOYMENT.md:62-70 (#1154's 43-minute measurement), handbook/batch-orchestration.md §measurements (#1140's 13-minute slot).
35. GitHub Actions cost, second pass — why the bill kept rising, and Depot assessed (researched 1 Oct 2026)
The finding: the bill is about $600 a month and rising, about 80% of it is the test suite on the two 8-core runners, and it rose because each suite run doubled in cost while the number of runs stayed flat. Every ship.yml and land.yml run from 17 Aug to 1 Oct was enumerated job by job through the API. Full write-up, method, decisions and the programme: outputs/docs/actions-cost-research-2026-10.md.
ISO week
36
37
38
39
40 (pace)
ship + land, est. $
71
68
45
85
~127
Why it rose: the suite got slower. Inside the land gate it went from 3.0 to 10.5 minutes on 8 cores between 21 Sep and 1 Oct (about 520 → 781 suites). 1.5× the suites for 3.5× the time points to a serial tail, not growth.
Why it rose: the 8-core move bought speed, not savings. The 20 Sep move (#2073, #2077) costs about the same per core as 2-core, and included minutes never apply to larger runners.
Why it rose: the suite runs too often. Each landed PR pays about 2.3 full suite runs (1 at ready, 0.54 repeats, about 0.8 at land); 87% of landings carry one PR; and 72% of landed PRs never touch the app, yet each runs all ~600 app suites.
Why 13 Sep's programme (§27) did not show. It cut week 38 to $45. The 8-core move and the suite's slowdown overtook it within a week, and nothing watched the cost of one suite run.
Depot: no. Its 8-core rate is $0.024/min against GitHub's $0.022. Per-second billing recovers only 8–13% here, so at equal speed it costs the same ($366 against $372 a month). Any saving rests on a speed-up nobody has measured on our suite. It also puts our secrets on a third party's machines, and its cache is not isolated by branch.
Cheaper runners exist (Ubicloud ~$0.005/min, RunsOn ~$0.003 spot, Hetzner AX42 ~£90 flat). Deferred: they only cut the rate on a runtime and run count that are both several times too high.
Decided (Dave, 1 Oct): the PR-side full suite runs once, at ready (again only after a red), with a targeted subset for non-app PRs.
Decided (Dave, 1 Oct): landing moves to four fixed windows a day, plus land-now, and a red batch is bisected. Two windows were costed at only $5–10 a month more saving; they are reviewed at the re-measure.
Decided (Dave, 1 Oct): the runner vendor waits for a two-week re-measure.
Estimated effect: about $600 → $90–160 a month (75–85%).
Filed as #2950 epic-actions-cost-2, children #2951–#2957, organised by landing unit. The open backlog that edits the same files is folded in: #2576, #2579, #2614, #2695, #2800, #2801, #2802, #2803, #2809, #2915, #2938, #2942, #2945, #2946.
Sources: Depot's pricing, runner-types, overview and security pages; GitHub Docs runner pricing, Actions billing and merge queue; the vendor pricing pages of Ubicloud, RunsOn, Namespace, WarpBuild, BuildJet and Blacksmith; Hetzner's 15 Jun 2026 price adjustment; the runs-on.com CPU benchmark (24 Sep 2026); the repo's own run history via the API. Links in the write-up.
Full write-up
From outputs/docs/actions-cost-research-2026-10.md
GitHub Actions cost, second pass — why the bill kept rising, and one combined programme (review, 1 Oct 2026)
Asked by Dave on 1 Oct 2026, in two steps. First: "would Depot reduce our Actions costs and the time runs take?" Then, once the answer was no: "widen this research to look at all options… I am seeing no reduction in costs whatsoever… I must see a vast reduction in the number of action minutes we're using." He then asked for the related open backlog to be folded in, so changes that each land alone go together and the process docs are edited once. Companion to outputs/docs/actions-cost-research-2026-09.md (the 13 Sep review and its baseline), which this document does not replace.
Status: decisions made by Dave in the session that wrote this (§ What was decided). Filed as epic #2950epic-actions-cost-2 with seven children, #2951–#2957 (§ The programme). No workflow has been changed by this review.
The short answer
The bill is about $600 a month and rising. About 80% of it is the test suite running on the two 8-core runners, and it rose because each suite run doubled in cost while the number of runs stayed the same.
The suite got 3.5× slower on the same box in ten days. Inside the land gate the suite took 3.0 minutes on 21 Sep and 10.5 minutes on 1 Oct on linux-8core. Over the same period it grew from about 520 to 781 test files: 1.5× the files for 3.5× the time. That shape points to a serial tail rather than to growth. The leading suspects are gitfixture1134.js, which re-runs 28 git-spawning suites one after another, and npm ci, which reinstalls node_modules from scratch on every shard and every gate.
Moving to 8-core larger runners (#2073, #2077, 20 Sep) bought speed, not savings. An 8-core minute ($0.022) costs about the same per core as a 2-core minute ($0.006), and included minutes never apply to larger runners. A land gate went from about $0.10 to $0.13 when it moved, and to about $0.25 once the suite slowed.
Every landed PR pays for about 2.3 full suite runs: 1 when it is marked ready, 0.54 repeats on later pushes (fixes, merge-main commits, re-readies) and about 0.8 in the land gate.
87% of landings carry one PR (285 of 325 successful land runs in 14 days). The sweep dispatches any approved PR 90 seconds after it is queued, and about 29 PRs land a day.
72% of landed PRs never touch the app (outputs/bakery-app-site/), yet every one of them runs all ~600 suites that boot or read the app bundle.
Depot would not help. Its 8-core runner costs $0.024/min against GitHub's $0.022. Per-second billing recovers only the 8–13% that GitHub's per-job rounding adds to our 8-core jobs, so at equal speed the two cost the same. Any saving depends on a speed-up that nobody has measured on our suite. No runner vendor is the main lever. Even the cheapest (Ubicloud, RunsOn, a rented server) only cuts the rate, and it would be applied to a runtime that is 3.5× too long and a run count about 5× higher than the process needs.
What cuts the bill: make the suite fast again, run it once per PR instead of 2.3 times, and land in batches at fixed times. Together these take an estimated $600 down to about $90–160 a month (75–85%), with a decision on a cheaper runner left until that is measured.
Why the 13 Sep programme did not show on the bill
It did work, for one week. The 13 Sep changes landed between 14 and 19 Sep: re-shard 6 → 3, the landing receipt fix, drafts skipping the suite, advisory jobs merged, ubuntu-slim. Week 38 came in at $45, down from $68–71. Then three things landed on top in the following ten days:
the move to 8-core runners (20 Sep);
the suite's slowdown from 3.0 to 10.5 minutes;
a higher landing rate (up to 52 PRs a day).
So week 39 was $85 and week 40 is running at about $127. The saving was real but was overtaken within a week. Nothing watched the cost of one suite run, so its doubling showed up only on the bill. Child A adds that watch.
The evidence
Method
Data. Every ship.yml run (4,772) and land.yml run (1,131) from 17 Aug to 1 Oct 10:00Z was listed through the API, and every job was fetched, including re-run attempts. For the 7 days to 1 Oct, every job of every workflow was enumerated: 8,132 runs, 6,056 jobs that actually ran.
Billing model. Each job that ran is billed ceil(completed − started) minutes at its runner's rate, the same rule as the 13 Sep review (GitHub's per-run billable data is still retired). These are estimates from timestamps, not an invoice; the 13 Sep calibration put the model within about ±10%.
Sampling.project-status.yml and feedback-status-sync.yml were scaled from 400-run samples. Everything else is a full count.
The trend (ship.yml + land.yml)
ISO week
Ship runs
Ship runs with a suite
Land runs
Est. $
34 (from 17 Aug)
233
(suite inside test)
0
7
35
660
(suite inside test)
70
32
36
665
281 (6 shards, 2-core)
327
71
37
571
390 (6 shards, 2-core)
279
68
38
676
279 (3 shards, 2-core, then 8-core)
139
45
39
1,349
362 (2 shards, 8-core)
218
85
40 (Mon–Thu 10:00)
618
157
98
62 (about $127/week pace)
The count of suite runs per week is flat at about 280–390. What moved is the cost of each one: the average 8-core suite run cost $0.099 in week 38, $0.131 in week 39 and $0.213 in week 40. The median land gate went from 4.9 to 8.5 minutes between weeks 39 and 40.
The last 7 days by runner (to 1 Oct 09:40Z)
Runner
Jobs
Billed min (rounded per job)
Exact min
Rate
Est. $
linux-8core-ship (suite shards)
604
2,526
2,243
$0.022
55.6
linux-8core (land gate)
197
1,423
1,318
$0.022
31.3
ubuntu-latest
1,597
1,784
978
$0.006
10.7
ubuntu-slim
3,658
5,144
3,261
$0.002
10.3
Total
6,056
10,877
7,800
~108
The 8-core runners are 81% of the money and 36% of the minutes. The ubuntu-slim jobs are the reverse: half the minutes and a tenth of the money. Over a third of the slim minutes billed are rounding (5,144 billed against 3,261 actually used), because a 15-second job bills a full minute.
Inside one suite run:
Ship shard (suite (1), suite (2)): median 3.6 minutes, p90 5.5. Checkout takes about 8 s; the rest is npm ci and the suite.
Land gate: median 6.6 minutes over the week, rising to about 10.5 minutes of suite alone by 1 Oct. 143 gate logs were read and none ran the suite more than once. The MAX_ATTEMPTS=3 retry loop exists but did not fire.
classify and journeys (every PR event, on ubuntu-slim): about 1 minute each, of which about 50 s is a fetch-depth: 0 clone of the 473 MB .git.
Per PR (the 14 days from 17 Sep)
Volume. 453 PR branches; 379 ran the suite at least once, 629 suite runs in all ($92). 409 distinct PRs landed, about 29 a day (range 9–52).
Suite runs per branch: median 1, p75 2, p90 3. One branch ran it 30 times: dependabot #2040, every run a rebase.
What triggered each suite run (inferred from each branch's sequence of head SHAs):
Trigger
Suite runs
$
First run, at ready
375
56.58
Repeat: a new commit (fix or dependabot rebase)
169
22.37
Repeat: a merge of main into the branch
72
11.66
Repeat: the same SHA again (reopen or re-ready)
13
1.63
Outcomes. 511 green, 97 red (89 with a red shard), 21 cancelled as superseded.
Not a cost driver: suite runs on PRs that were later abandoned came to about $1.
Landings (the same 14 days)
Runs. 378 land runs: 325 green, 50 failed, 3 cancelled. The failures are cheap ($4.21); most stop in the precheck inside a minute.
PRs per successful run: 1 PR in 285 runs, 2 in 20, 3 in 8, 4 or more in 12.
Why batches rarely form.land-sweep.yml dispatches a PR alone once its queued label is 90 seconds old (MIN_LABEL_AGE_S). Dave's land-approved + /go arrive spread through the day, so there is rarely a second PR waiting.
Forced solo landings (#2922, landed 1 Oct): a batch may not include a PR that edits the landing pipeline, and may carry at most three review-sensitive PRs.
Pushes to main and the small workflows
Pushes to main. 366 ship.yml push runs a week cost about $2.34; none ran the suite. 168 of them are the robot's LIVE-VERSIONS: refreshed after this staging deploy commit, one per staging deploy.
Issue-driven workflows.project-status.yml (1,838 runs a week from issue events), land-sweep.yml (617), feedback-status-sync.yml (3,266, 99% skipped), epic-children-check.yml and release-train.yml together cost about $2.65 a week.
Crons. About 550 scheduled runs a week, nearly all on ubuntu-slim.
These are large in run count, small in money. They are child E, not the headline.
The suite itself
Size.LIST_ONLY=1 run.sh lists 780 suites, not the ~515 that the comments in ship.yml:599-607 and run.sh:11-19 still describe.
What they read. About 600 suites read the generated app bundle (4.2 MB); 445 boot it in jsdom; 28 use pglite; 28 drive real git fixtures. None drives a real browser.
No timings are kept.run.sh prints no per-suite timings, and the raw 5 Sep timing file was never committed, so there is no way to see from the repo which suites caused the slowdown.
Targeted selection is feasible. About 320 suites reference nothing outside outputs/bakery-app-site/; for a PR that does not touch that folder they cannot change their verdict. The repo already has the fail-closed tool for this: run.sh's $READS selection, used by the site-only gate (#1957).
Depot, assessed on this data
Depot's prices (runner-types page, 1 Oct 2026):
Linux 2 vCPU $0.006/min · 8 vCPU $0.024/min · 16 vCPU $0.048/min, billed per second.
No 1-vCPU size.
Plans: Developer $20/month with 2,000 included minutes; Startup $200/month with 20,000. Larger runners use up the included minutes at their size multiplier.
Org-owned repos only (we qualify since #1724). Runners are ephemeral EC2 in AWS us-east-1.
The 8-core suite work, at equal speed: $366/month on Depot against $372 on GitHub. Per-second billing saves about 10% on our 8-core jobs; the higher rate takes about 9% back.
The saving depends on an unmeasured speed-up:
Depot speed-up on our suite
Monthly saving
1.0×
about $0
1.3×
about $80
1.5×
about $120
2×
about $180
The "up to 3× faster" claim is from Depot's own BuildKit benchmark. The only independent CPU benchmark (runs-on.com, 24 Sep 2026) does not include Depot.
The small jobs would get dearer.ubuntu-slim jobs would cost about 2× more, because Depot's smallest runner is 2 vCPU.
Risks: our secrets (GH_MERGE_TOKEN, the Cloudflare deploy token) on a third party's machines; no public SOC 2 Type II report found; a cache that is not isolated by branch; a queue-forever stall in land-main if Depot is down; separate billing outside the GitHub budget alert; and GitHub's postponed $0.002/min self-hosted platform charge, which would apply to every Depot minute if it returns.
Verdict: no.
Every option considered
Option
8 vCPU $/min
Est. monthly for ~15,000 8-core min
Notes
GitHub larger runner (today)
0.022
$330
Rounded per job; included minutes never apply
Depot
0.024
~$360
See above
Blacksmith / BuildJet / WarpBuild
~0.016
~$230–240
Blacksmith's 8-vCPU rate is from secondary sources only
Namespace
0.008 prepaid / 0.012 overage
~$130
$100/month team plan
Ubicloud
0.005 standard / 0.008 premium
~$75–120
Label change plus a GitHub app
RunsOn (our own AWS)
~0.0027 spot / 0.0071 on-demand
~$65–130 incl. €300/yr licence
Needs an AWS account
Hetzner AX42, self-hosted
flat
~£90 flat (€97.30 + €49 setup, ex VAT)
Prices more than doubled on 15 Jun 2026; we own patching, uptime and runner security
Dave's Mac as a runner
flat
£0
Rejected: personal machine, sleeps, macOS not Linux, holds credentials
GitHub merge queue
n/a
n/a
Appears unavailable for private repos on Team (docs and a Jul 2026 community request; unverified)
Local sign-off (gh-signoff)
n/a
n/a
An honour system; land.yml stays the authoritative gate either way
Process: fast suite, once per PR, batched landings
as today
~$40–95 for the 8-core share
Chosen
A cheaper runner would multiply whatever the process changes leave, so it is decided after the re-measure (child F), when the remaining suite minutes are known. GitHub's self-hosted platform charge ($0.002/min, announced 16 Dec 2025 for 1 Mar 2026) was postponed, not cancelled. If it returns it adds about $30/month to any self-hosted or vendor-managed option at today's volume.
What was decided (Dave, 1 Oct 2026, in the session that wrote this)
PR side: the full suite runs once, at ready.
It runs again only if the last suite run on that PR was red.
A PR that touches nothing in the app, its tests' shared libraries, the lockfiles, the gate or the workflows runs a targeted subset: its changed suites plus the suites that reference the folders it touched. This is a fail-closed grep, the same shape as #1957's $READS.
Sessions run the suite locally (free) before marking ready.
land.yml always runs the full suite on the merged tree, so the safety net is unchanged.
This reverses part of 13 Sep's B2 decline: the PR-side suite stays, but only once per PR.
Landing: four fixed windows a day.
About 09:00, 13:00, 17:00 and 21:00 UK; a window with nothing queued runs nothing.
A land-now label is Dave's immediate path. Hotfix and release-fix keep their exemptions.
A red batch is split in halves to find the culprit, instead of landing every PR alone.
Two windows (09:00 and 21:00) were costed and deferred. Estimated gates per day are about 5.2 for two windows against 6.6 for four, once the bisection of bigger (about 15-PR) batches is counted. That assumes 2.5% of ready-green PRs still go red on the combined tree, which is not yet measured. The difference is about $5–10 a month, for up to 12 hours of staging lag. Start at four and review at the re-measure, with the real red-batch rate. The window times live in one config list, so changing to two is a one-line edit.
Runner vendor: decided after a two-week re-measure (child F).
Expected effect, with the arithmetic
The starting point is the 8-core spend at the week-40 pace, about $100 a week, split 64% PR side and 36% land, as in the 14-day split ($92 / $51).
Lever
Factor
Basis
Once at ready
PR-side runs × 0.6
375 first runs of 629 in 14 days
Targeted subset for non-app PRs
× ~0.39 full-suite equivalents
0.28 app PRs × 1 + 0.72 × ~0.15
Land windows with bisection
gates × ~0.2
~6.6 gates/day against ~27
Suite back to ≤ 4 minutes
per run × ~0.4
10.5 → ≤ 4 min
8-core share: $100 × (0.64 × 0.6 × 0.39 + 0.36 × 0.2) = ~$22/week without the speed fix, ~$9/week with it. That is $40–95 a month.
Small jobs: about $20 a week now, about $12–15 after child E.
Total: about $90–160 a month against about $600, a 75–85% reduction. Billed minutes fall by a similar order, and fastest of all on the 8-core runners.
These are estimates. Child F measures them.
The programme: #2950 epic-actions-cost-2, organised by landing unit
Each child is one PR and one landing. A PR that edits ship.yml, land.yml, land-sweep.yml, gate.sh, .github/actions/gate/ or security-gate.* lands alone (#2922), so the open issues that edit the same files were folded into the child that edits them. Each moved issue becomes a sub-issue closed by that child's PR. Nothing was closed by this review. Mechanism docs travel with their mechanism; process-only docs go in D.
Child
What
Absorbs
Lands
Sequence
A. #2951 suite-speed-regression
Per-suite timings in run.sh; root-cause 3.0 → 10.5 min; fix the gitfixture1134.js serial tail; cache outputs/tests/node_modules keyed on the lockfile; an advisory suite-time budget, proven by planting a slow suite (#22); correct the stale suite counts
— (nothing open covers it); include the #2697/#2834 ESLint suite in the baseline
alone (gate.sh, gate action)
First; after dependabot PR #2927 (tests lockfile). Before the Mon 5 Oct 12:00 freeze if possible
B. #2952 ship-yml-suite-once-at-ready-and-targeted
Once-at-ready and targeted selection in classify, with plants; blobless clone for classify/journeys
#2803 (stop the LIVE-VERSIONS robot commit), the ship.yml half of #2938, #2915's ship.yml rows, the ship.yml part of stale draft PR #2583 (#2579)
alone
After A; #2804, #2805, #2807 (#2799) follow on top
C1. #2953 security-gate-batch-fixes
Batch-classification fixes so phantom hits stop forcing split runs
#2946, #2942, #2614 (salvage stale draft PR #2660, which predates #2926)
alone, security-reviewed
After the 5–6 Oct release
C2. #2954 land-windows-and-bisect
Four windows from one config list, land-now, a higher BATCH_CAP, bisect on red, per-window telemetry (batch size, red or green, bisect gates); docs in WAYS-OF-WORKING §3.2/§3.2a, handbook/shipping.md, batch-orchestration.md, .github/DEPLOYMENT.md
#2576 (salvage stale draft PR #2589), #2945, the land.yml half of #2938, #2801, #2915's land-sweep.yml rows
alone
After C1; #2940 (#2871) rebases on it
D. #2955 ci-process-docs
No merge-main pushes to a ready PR; fix batch.md:491 against :912; local suite before ready; group issues into one PR; AGENTS.md #1 wording
#2579 (docs parts of stale drafts #2581/#2583), #2800, #2809
docs path
Any time; written against B's and C2's specs
E. #2956 small-workflow-trim
Debounce project-status.yml; the rest of #2915; check #2754
#2915 (remaining rows), #2802
normal
Coordinate with draft PR #2825
F. #2957 actions-cost-remeasure-and-decisions
Two weeks after A, B and C2: re-measure by this method and bring Dave the runner vendor, 4 → 2 windows and #2695 (CodeQL cost) as one budget review
#2695 (decision only)
—
Last
Risks to name on C2. #2604 (a separate GitHub identity for sessions): wrong-actor refusals already broke a six-PR batch (#2945). Batching only pays if a batch is not refused for reasons unrelated to its code. #2611 (least-privilege admin separation) may change the landing identity.
What is deliberately left alone:
land.yml re-running the full suite on the merged tree: it is the only gate that decides anything.
The docs-only and site-only classifier as the only allowlist: still no trigger-level paths-ignore (#1407).
The test job's exact name and the rule that it runs on a red shard.
cancel-in-progress on PR runs.
How to re-measure (child F, #2957)
Use the 13 Sep method plus this one:
Page list_workflow_runs for the window and fetch every job with filter=all.
Bill each job that ran at ceil(minutes) × its runner's rate.
Report per runner, per workflow, and per PR: suite runs per landed PR, and PRs per land run.
Read the suite's own timings (child A).
Targets:
a land-gate suite at ≤ 4 minutes median over 20 runs;
≤ 1.1 PR-side full-suite runs per landed PR;
≥ 70% of land runs carrying two or more PRs;
8-core spend under $25 a week.
If the 8-core spend is still material after that, put the Ubicloud / RunsOn / self-hosted question to Dave with the remaining minutes.
The repo's own run history through the API (window and method above); outputs/docs/actions-cost-research-2026-09.md; outputs/docs/batch-process-time-review-2026-09.md.
Notes
Last updated 27 Jul 2026
6. Architecture facts that constrain the business
The app is single-tenant.business_settings is one row (id=1, singleton constraint) and every RLS policy is authenticated full access — any signed-in user can read and write all data. One bakery per deployment. Nothing can be sold until this changes (#48). > Corrected 27 Jul 2026. This was true when written and has been wrong since v0.14.0, which > shipped #48. The app is multi-tenant: businesses + business_members + platform_admins, > business_id denormalised onto all 16 tenant tables so every policy is a flat > business_id = my_business_id(), self-serve sign-up via the create_my_business() SECURITY > DEFINER rpc, and isolation proven with a real throwaway tenant (it saw 1 ingredient, not our > 76; cross-tenant insert/update/delete all blocked). The anon branding policy was dropped > deliberately. The commercial blocker is no longer architectural - it is billing, staging > (#137), the legal stack and support tooling. Anything still citing "#48 blocks selling" is stale; > this paragraph misled the 27 Jul commercial report until it was corrected the same day.
Deployment is a single-file index.html app served as Cloudflare Workers static assets + Supabase (Postgres, auth, edge functions). Cheap to run per tenant; no build step. (Corrected 4 Aug 2026: this said "on Netlify" — serving moved to Cloudflare Workers in v0.85.0/v0.85.1, #251/D-20, and Netlify is retired.)
Supermarket scraping is a dead end — the big sites block automated reads. Tried and removed in v4.5. Don't revisit.
Open Food Facts was the only realistic external data source: free, barcode-addressable, and it publishes its own allergens_tags/traces_tags (a genuinely independent second opinion — #45). Caveats: crowd-sourced, patchy for UK own-brands, not an authority. Corrected 28 Jul 2026 — we removed it entirely (v0.75.0), and the caveat turned out to be the whole story. "Patchy for UK own-brands" understated it: own-brands are roughly half a home baker's shopping, and their ingredient data is the retailer's own commercial asset, so it is not in any open database and will not appear in one. A source that misses half the catalogue cannot be the primary route to a declaration, and once #249 (photograph the pack, read it with Claude vision) covered 100% of packs including own-brands, OFF's remaining value — the independent allergen second opinion — was not worth three UI surfaces, a barcode field, a camera scanner and a public-API dependency. The finding that survives is the negative one: there is no external database route to UK own-brand declarations. The pack itself is the only source that always exists.
Notes
Last updated 17 Jul 2026
6a. Superuser / vendor access — decide before building (#57)
Deferred to Later on 17 Jul 2026 (first outside bakery is months away), but the thinking is worth keeping, because the obvious implementation is the dangerous one.
Some of it already exists.platform_admins (v0.14.0) gates the internal backlog and debug logs, and lets an admin list every business. Missing: a console, and "act as" a bakery.
⚠️ Do not add or is_platform_admin() to every tenant policy. It's the one-line version and it would undo the isolation proven in v0.14.0: one compromised admin account would reach every customer's data, and the promise degrades from "your data is yours" to "yours, plus whoever we trust". Prefer time-boxed, explicitly logged impersonation, or a service-role console kept outside the customer app.
A bakery's recipes are its IP. Silent vendor access is a product decision, not a technical one. Support access should be logged — ideally visible to the bakery. Worth deciding before a customer asks how we handle it, not after.
Timing argument: with zero other bakeries, the Supabase dashboard is the admin console, and a multi-bakery UI would be designed against imagined needs. The day customer #1 says "my labels look wrong", it becomes urgent — so it should land before the first customer, not during their first problem.
Notes
Last updated 10 Aug 2026
7. Open questions
Market size — how many registered home food businesses in the UK?Bounded 31 Jul 2026 (§12): TAM/SAM/SOM now modelled with stated assumptions. What stays open is narrower: the official PPDS in-scope count (read the 2019 Impact Assessment, ukia/2019/144, from a browser) and the true active home-based stock (FHRS API business-type query or an FSA FOI).
Would a baker pay? Nothing here is validated with a real customer. The cheapest falsification of every strategic claim above is two home bakers who aren't us.
PAL propagation — is auto-copying a supplier's "may contain" defensible, or does it need to become an explicit per-ingredient risk-assessment decision?
Nutrition exemption — confirm the micro-enterprise position before building #28.
Pricing — £19–55/mo is the anchor(corrected 10 Aug 2026: FoodCore repriced — the anchor is now £25–65/mo inc. VAT, see §3 and the master competitor sheet). Where do we sit, and is the moat worth a premium?
Who verifies? If a baker's staff verify allergens, we need roles (#38) sooner than ranked.
Notes
Last updated 19 Jul 2026
9. Terminology — the "ingredient hierarchy" (researched 19 Jul 2026)
The app has several tiers that are all "ingredients" at heart but mean different things, and they needed distinct names. Decision (Dave, 19 Jul): Ingredient / Supplier product / Declaration / Sub-recipe. Full design + the string-by-string relabel plan in ingredient-terminology-design.md. The evidence:
Ingredient = the generic tier. UK labelling law (retained Reg. 1169/2011 / FIC) defines an ingredient as a substance used to make a food and present in the finished product; a compound ingredient is one made of several ingredients, whose parts are sub-ingredients shown in brackets. On our finished label the entry is the generic name ("Caster Sugar"), so the law reserves "ingredient" for the generic tier and "sub-ingredients/declaration" for what's inside a bought compound item. Compliance-first ⇒ we follow it.
The purchased tier is a "product/item", not an "ingredient". Product-catalog / PIM modelling universally uses Product (abstract) → Variant → SKU; USDA FoodData Central splits Foundation Foods (generic) from Branded Foods (brand-specific); recipe/food-cost tools link each ingredient to a supplier item. All three put a product/item word on the buyable brand/pack/vendor thing.
Why "Supplier product" not "Product" or "Pack". "Product" is already used in the app for sellable finished goods (products.unit_price) — a collision. "Pack" breaks for loose/bulk items. "Supplier product" is unambiguous and the "Supplier" qualifier keeps it distinct from sellable Products. Rule: never shorten it to "product" in the ingredients area.
The sellable output stays "Product". In the standard chain — raw materials/ingredients → recipe (BOM) → finished goods, with menu items the POS word — our products (the bake output with a price) are the finished goods. We considered renaming them to "Finished good" (BOM-standard, and it would free "Product" for Tier 2) or "Menu item" (POS-standard), but decided to keep "Product" (Dave, 19 Jul): "Finished good" reads too industrial for a home baker and "Menu item" assumes a fixed menu they may not keep. So the chain is Ingredient → Supplier product (buy side) and Recipe → Product (sell side); the "Supplier" qualifier is what permanently separates the two.
**The relabel is an inversion, so the end goal is a real rename.** Because "ingredient" currently sits on the purchased tier in the DB but will mean the generic tier in the UI, a permanent display-alias would mislead future maintainers (and Claude). The plan is therefore staged: relabel words now (+ a NAMING-MAP comment), and do a full front-and-back table/column rename bundled with the later Ingredient-first restructure, behind backward-compatible views so the live app never sees code and schema disagree.
Tables don't change — this is a display-layer relabel (label_names→"Ingredient", ingredients→"Supplier product"), like feature_requests→"Backlog". The current app has "ingredient" on the purchased tier, i.e. backwards; we flip the words, not the schema. (Superseded: the full DB rename this section deferred landed in v0.39.2 — tables now carry the user words, and neither label_names nor feature_requests exists any more.)
Sources: retained EU Reg. 1169/2011 (legislation.gov.uk); FSA packaging & labelling guidance; commercetools / Elastic product-variant-SKU modelling; USDA FoodData Central (Foundation vs Branded Foods); recipe-costing software docs (Reciprofity, Supy); food-ERP BOM / finished-goods glossaries (Folio3, Wherefour, Toast). Links in §Sources.
Notes
Last updated 3 Aug 2026
9a. Is "Recipe" a barrier for non-bakers? No — keep it (researched 3 Aug 2026)
Raised by Dave alongside the #243 bake→batch rename: "Bake" was baker-only, so is "Recipe" the same class of problem for the widened D-24 market (butcher, deli, caterer, jam maker, farm shop)?
Finding: no. "Recipe" is trade-neutral in exactly the way "Bake" was not. Three lines of evidence, all checked 3 Aug 2026:
The regulator uses it for non-bakery trades. GOV.UK publishes PPDS allergen-labelling guidance specifically for butchers (sausages, marinated steaks, burgers), and Food Standards Scotland's Butchers guide (Apr 2024) frames ingredient/allergen control around the recipes for made-on-site products. A butcher's sausage has a recipe in the regulator's own vocabulary.
Every comparable UK product uses it as the category word. Kafoodle ("recipe management"), MenuIQ ("recipe and menu management software"), Nutritics, and FoodCore — the named competitor in §11 — all describe the same object as a recipe. A user arriving from any of them, or from a Google search, is searching the word "recipe".
The alternatives are worse for our segment. "Formulation" and "Specification/spec" are the manufacturing/BRC words and read industrial to a micro business (the same reasoning that kept "Product" over "Finished good" in §9, Dave's 19 Jul decision). "Method" describes the steps, not the costed ingredient list the app actually models.
The real D-24 risk is the examples, not the noun: recipe copy and placeholders should keep rotating non-bakery examples (a sausage rub, a chutney batch, a marinade) — which is already the copy policy. Decision recorded: keep "Recipe"; no rename planned. If future beta feedback from a non-baker contradicts this, correct this entry rather than deleting it.
Sources: GOV.UK "Prepacked for direct sale (PPDS) allergen labelling changes for butchers"; FSA allergen guidance for food businesses (food.gov.uk); Food Standards Scotland "Food Standards Guide: Butchers" (Apr 2024); kafoodle.com; menuiq.co.uk; nutritics.com; foodcore.io.
Compared our architecture against Stephen G. Pope's "Build A SaaS Startup With Claude Code (22-Min Crash Course)" (YouTube, s47C9_qDvJs) and its linked resources. Full report: video-methodology-comparison.md. Headlines: the video's stack (Next.js + shadcn, Supabase, Claude Code + Supabase MCP + CLAUDE.md, GitHub, Vercel push-to-deploy) validates our Supabase + Claude choices; we're ahead on production discipline (tests, migrations + rollbacks, RLS multi-tenancy, drift check, post-deploy verification — the video shows none). Real gaps on our side: no git history and no push-to-deploy CI (→ backlog #136), and UAT shares the prod database (→ backlog #137). Decision: no Next.js rewrite pre-revenue; modularise the single file incrementally; revisit a framework only if SEO/SSR demands it at commercialisation.
15. "Data quality" naming — what the food industry and data standards actually call this (researched 15 Aug 2026)
Context: user feedback said the Data quality screen's category names (Blocked / Quietly wrong / Tidying) confuse, and Dave asked whether the screen's own name has a more food-industry-native alternative, and whether governing bodies have standard vocabularies for classifying data issues. Findings from a web pass (all sources below, dated within months except where noted):
The food-safety audit world classifies findings by SEVERITY — critical / major / minor — and this is the vocabulary our audience already meets. BRCGS (Global Standard for Food Safety, §2.3.1) grades every audit non-conformity: critical = a direct food-safety or legal issue; major = significant doubt about product conformity, or a legal/safety issue if no action is taken; minor = a partial miss that does not affect safety, legality or quality. That maps 1:1 onto our three classes (Blocked → critical, Quietly wrong → major, Tidying → minor). SALSA — the UK scheme aimed at exactly our micro/small-producer segment — uses the same non-conformance framing. Caveat for UI use: in audits these words grade food-safety non-conformities; borrowing them for data findings risks implying the app audits food safety.
The FSA's own small-business register is deliberately jargon-free: SFBB talks in "checks" (opening/closing), "problems" (the diary's problem log records what went wrong and the action taken), and the "regular review". No severity taxonomy at all — the closest SFBB concept to our screen is the regular review ("look at problems and patterns"). "Checks" is therefore both the most natural word for our audience and a collision risk with SFBB's physical daily checks (#36).
The data-management standards classify by DIMENSION, not severity — wrong axis for users. DAMA-DMBOK's canonical six (accuracy, completeness, consistency, timeliness, uniqueness, validity), ISO/IEC 25012 (defines the characteristics) + ISO 8000 (how to verify/exchange conforming data). Useful internally as a checklist of what our findings sweep should cover (we do completeness + consistency today; timeliness partially via stale-declaration findings); useless as user-facing names for a small food business. GS1's data quality programmes (attribute accuracy vs the physical product, e.g. declared net content) are the closest food overlap but are built for GDSN/large-retail data pools, not PPDS microbusinesses.
Net: there is no governing-body standard for naming in-app data issues, but there IS a strong industry convention for grading finding severity (critical/major/minor), and the FSA sets the plain-English register. Candidate screen names discussed: "Health check" (common in small-business software, no SFBB collision), "Checks", "Readiness"/"Label readiness", "Self-audit". Decided same day (Dave, 15 Aug 2026, #540): screen renamed "Health check", categories "Critical / Major / Minor" — implemented as v0.195.0 on the claude/data-quality-categories-g9jkv7 branch (the "every serious defect" sentence had already gone in v0.194.1).
16. Social monitoring — what the platforms actually allow (#579, researched 20 Aug 2026)
⚠️ Corrected 16 Sep 2026 (decision 1939 / #1940): the monitoring agent is retired. The findings below stay as the record of why the anonymous public-group route was chosen and why Facebook groups cannot be automated; they are not a live capability. Metricool posting of our own four channels is unchanged.
Research pass for the #579 monitoring agent (find conversations → draft replies → Dave-gated posting). The findings that will outlive the build:
Facebook Groups API is gone (22 Apr 2024) and nothing replaced it — Zapier, Make and Zoho all discontinued their Groups integrations. Every surviving "monitor Facebook groups" product (Devi AI is the category leader, ~$19–50/mo) is a browser extension driving the user's own logged-in session — exactly what our standing rule bans ("no group auto-poster extensions, ever"; the account at risk is Dave's irreplaceable personal profile).
Public groups were readable anonymously through third-party no-login scraping services, with the risk sitting on the vendor's infrastructure. That was the Phase 1 discovery route; it was retired with the agent and the vendor account is closed (#2988). Of the #218 register, rows 11–12 are confirmed public (Cake Sheds and Honesty Boxes UK 10.8K; UK Baking Community 2.8K); most others unverified, Finch Bakery confirmed private.
Native FB notifications are not a dependable feed: a member can set a group to "All posts", but Facebook still algorithmically decides which posts actually notify. Supplement for 2–3 priority groups, not a foundation. Real keyword alerts inside a group are admin-only — which is why the admin-partnership route (§12) doubles as the only fully clean automated monitoring of a private group.
Instagram Graph API (business accounts only): hashtag search capped at 30 unique hashtags per rolling 7 days (200 req/hr); polling media/comments that @mention or tag the account; API replies allowed ONLY to @mentions and comments on own media. No general keyword search; no commenting on strangers' posts. A Development-Mode Meta app suffices when serving only our own account (no App Review).
TikTok has no commercial discovery API: the Research API (keyword video search) is restricted to academic/non-profit researchers; commercial applications are rejected. Dave dropped TikTok from the monitoring exercise on this basis (20 Aug 2026). Own-account comments remain reachable later via Metricool's paid Inbox.
Metricool's MCP server (available to Claude sessions — and, verified 15 Sep 2026, working on the Free plan for reads and plain creates despite Metricool's page saying Advanced; only the review-mode create is 403 — see §28) covers scheduling, analytics, best-time and competitor data — no social listening, no inbox on the MCP. Paid listening tools (Brand24 ~$199/mo, Mention ~$41/mo, Awario $29–74/mo) cannot see inside Facebook groups either; deferred post-launch.
18. Beta-tester community — decided and stood up (#708, closed 28 Aug 2026)
Full reasoning: the Beta community plan (Google Drive since #2809 — outputs/docs/drive-index.md). Recommendation was "rent the peer-support half (a private Facebook Group), defer the ideas board" — see that doc for the two live measurements (population size, zero feedback_items usage) it was built on.
Two of the plan's pieces shipped as their own issues: #716 (read-only in-app "What we're working on" roadmap card — stage 1) and #761 (screenshot attachments + a global "give feedback" entry point, the fix for "Give feedback is buried"). Re-measured on 28 Aug, two days after #761 shipped: feedback_items is still 0 rows — real evidence against "it was just buried," not just the original 25 Aug guess.
The Facebook Group is live: "ProvenBatch Beta", created 28 Aug 2026, administered by the ProvenBatch Page (not Dave's personal profile), Private + Hidden. The plan's own text had gated group creation on "the 3rd tester lands" (~23 Nov) — Dave chose to start it early rather than wait, once business_settings showed betatester 2 / trialtester 2 / customer 2 (up from 1/1/2 on 25 Aug). Logged as a deliberate deviation from the plan's dated gate, not a silent skip of it.
Still open, deliberately: whether tester-initiated threads justify keeping the group past the beta, and the #708-stage-2 shared-ideas-board go/no-go — both need real usage evidence the group doesn't have yet. Tracked as a follow-up issue rather than left inside #708, refs #708.
Source: production business_settings / feedback_items, read 28 Aug 2026; #708, #716, #761.
Notes
Last updated 30 Aug 2026
19. Short-form content craft, trend automation, AI video tooling and the mailing list (researched 30 Aug 2026)
Four-agent web research pass for the content-upgrade session (epic #181). What the social batch and the video pipeline are now built to; the numbers moved into outputs/video/README.md § Pacing and safe zones the same day.
19a. Short-form craft — the numbers
The first ~1.5 seconds decide distribution on TikTok; ~90% of underperforming videos fail in the first 3 seconds. The best-tested hook archetype across a 34,635-clip dataset is show the outcome first (the finished label printing — never a logo or scene-setting).
Visual change every 1.5–2 seconds is the 2026 pacing norm for sub-60s video (below ~1.2s reads as noise). Our explainer cards previously held 3.0–4.6s — exactly the "screens don't change fast enough" complaint; re-paced 30 Aug (18–26s totals, was 26–34s).
Length: TikTok completion optimum 21–34s; Reels 15–30s; Shorts 30–45s; product demos with 2–3 steps 30–60s. Completion, not length, is the metric the algorithms reward.
~85% watch muted (still the standard citation); burned captions add 12–15% completion. Our beat cards ARE the captions, which the re-pace preserves.
Safe zone: a centred 900×1400 of the 1080×1920 canvas clears every platform's UI (TikTok bottom ~270px + right ~100px, wider since Jan 2026; Reels top ~200px + bottom ~400px). Covers: 3–6 bold words, centred — the grid crops top/bottom and renders ~120px wide.
B2B/SaaS specifics: micro-demos (one feature, one clip) beat full demos; sources split on UGC-style face-to-camera vs silent screencap — the reconciliation is face/hook for the problem, real UI for the proof. For our audience the UK small-food-business precedent is strong (1.5m UK companies on TikTok; butchers and bakers among the documented successes with behind-the-scenes content).
Pre-launch playbook: start the waitlist early, 2 posts/week beats a launch-week blitz, and CTA-ladder — teaser posts ask for a follow/save; only posts that showed a real workflow ask for the waitlist/beta click. Email the list every 1–2 weeks.
19b. Trend discovery and automation — the constraints that decide the design
No official trends API exists on any of our four platforms. TikTok Creative Center is a logged-in browse (its no-login access looks stale as of 30 Aug); Instagram's official Trending Audio surface is US-only for pro accounts; YouTube's Trending page was retired July 2025 in favour of in-Studio Trends/Inspiration. Third-party "trend APIs" are scraper-style — ToS-grey, breakage-prone, and not adopted (§16's conclusions stand; the scraper vendor itself was retired from ops 16 Sep 2026, decision 1939 / #1940).
Trending audio cannot be scheduled: a sound must be attached in-app for the post to join the sound's discovery page, and Business IG accounts get a restricted (commercially-cleared) music library. So trend-audio posts are Dave-from-the-phone by construction.
**The enforcement line platforms actually police is automated engagement, not automated posting.** Scheduling via an official API partner (Metricool) is safe; auto-follows, bot comments and DMs are where accounts die. Never automate replies.
Consequence (Dave, 30 Aug 2026): trend intake is a weekly, draft-only step of the metricool-drafting routine — at most 2 [trend] drafts a week from fetchable sources (Metricool/Buffer/Later trend roundups, the UK seasonal food calendar, fsa-watch/recall-watch output), all via createScheduledPostForReview. Nothing publishes itself.
19c. AI-generated video — the rules that outlived the tool
Corrected 1 Oct 2026 (#2988): the generation vendor evaluated here (30 Aug – 1 Sep 2026) is closed, so its product, pricing and API notes are removed. The vendor-neutral findings stay.
Honest read for a software product: generated video is a b-roll and hook factory, not a demo generator.
Never put app UI through a generator. Generated or restyled UI mangles text, and for a labelling product the text IS the product. Screen recordings stay real (outputs/video/ pipeline); generated footage touches only the 0–3s hook and the connective tissue.
Highest-ROI uses for us: food-business b-roll (kitchens, label printers, hands, packaging — replaces stock footage; the models are strongest on slow, quiet motion), 2–5s openers/stingers, and layered poster output for covers/carousels. An avatar presenter is possible but must be disclosed and never posed as a customer (ASA misleading-testimonial territory).
Hosted outputs expire, so anything adopted must archive immediately to media.provenbatch.co.uk.
Disclosure: generators that sign C2PA metadata are auto-labelled as AI by TikTok/Meta/YouTube regardless — we declare proactively via Metricool's isAigc/isAiGenerated/ isAiGeneratedContent flags (rule now in social-content.md §2 and the metricool-drafting skill). EU AI Act Art. 50 transparency obligations apply from 2 Aug 2026; UK exposure is ASA/CAP misleadingness, sharpest for fake-customer avatars.
Findings from the #181 build-out (1 Sep 2026; method: outputs/docs/social-video-production.md): (a) TikTok's Commercial Music Library is where trending tracks come from, attached at DIRECT_POST publish; TikTok drafts drop music, so "preview with the real sound" doesn't exist off-platform. (b) Instagram business accounts get a restricted music catalog — mainstream trending songs are mostly absent; Metricool can attach what IS there at schedule time via audioConfiguration, and its search resolver silently binds a wrong track on fuzzy matches (measured — always verify the resolved title). (c) CapCut has no official API or MCP server (checked 1 Sep 2026): community "CapCut MCP" projects only write local draft files for the desktop app — timeline assembly we already do free in the sandbox; the parts we'd want (template library, animated captions, licensed-audio preview) are app-only. Revisit only if an official API ships. Sources: usecarly.com/blog/capcut-mcp, samautomation.work/capcut-api, github.com/sun-guannan/VectCutAPI (all read 1 Sep 2026).
19d. Mailing list — PECR reality and the ESP decision
A pre-launch waitlist does NOT qualify for PECR's soft opt-in (that needs collection during a sale/negotiation of a sale). Explicit consent is required — which the /beta form's separate, unticked updates_opt_in box already provides. Many of our prospects are sole traders = individual subscribers under PECR, so the strict rule applies to the core market.
Now (Dave, 30 Aug 2026): build on beta_requests — an updates-only path on /beta, updates excluded from invitation sending, HMAC-token unsubscribe preserving the no-read-path posture, campaigns as beta_emails kinds via the existing Resend rails. Issues filed from this session.
Later, when the list outgrows the rails (~250+ subscribers or GA): MailerLite — the one candidate with EU data residency and a DPA on every account by default (the right look for a compliance product), built-in double opt-in and GDPR forms, free to 250 subscribers / 2,500 emails a month, then from $12/mo. Buttondown is the API-first runner-up (100 subs free, ~$9/mo); Kit's free tier is big but cliffs at $39/mo; Mailchimp's free tier is no longer competitive. Becoming a sub-processor is a D-12 disclosure change — part of the adoption cost, not a surprise. Dated addendum in outputs/docs/gtm-channels-research.md.
This entry recorded a billing-API finding for a speech-to-text vendor that left the stack with transcribe (#1939) and whose account is now closed (#2988). Removed rather than kept as history. Section numbers after it are unchanged.
Notes
Last updated 7 Sep 2026
22. BabyLoveGrowth as a growth tactic — not worth buying, and not worth cloning (#1263, researched 7 Sep 2026)
Dave flagged babylovegrowth.ai via his requirements list: how much of their approach could be replicated in-house. They are a growth-tactics vendor, not a compliance-software competitor — they do not belong in §3. Verified against their live homepage, pricing page and docs on 7 Sep 2026, not from memory or a review-blog summary.
22a. What they actually sell
An AI SEO / GEO autopilot. Connect a site; they research keywords, write articles in the brand voice, auto-publish them, trade backlinks through a partner network, run a technical/GEO audit, and run Reddit/Quora agents that look for posts LLMs cite so the brand can be mentioned there.
$99/mo (struck £247 on the page — a display conversion, not a UK plan)
30 AEO/SEO articles/month; auto-publish to WordPress, Shopify, Wix, Webflow "and 8 more"; "$800/mo in backlinks from 4,000+ partner sites"; AI-visibility tracking across ChatGPT / Perplexity / Gemini (10 prompts × 3 models); site audit; Reddit & Quora agents; automated keyword clustering
Scale
$299/mo (struck £599)
Grow plus a dedicated SEO/GEO specialist, 2 hours/month, three extra languages, beta access, same-day support
Agency / white-label
$99/site/mo, plus $2,000 one-time for the branded dashboard
Resell the same machine. Not relevant to us.
A competing review (Distribb, updated 6 Aug 2026) quotes the same $99 / $299 shape and a 3-day trial. We did not take their "RebelGrowth is better" conclusion; the price and the product shape match the live page.
CMS list on the live page: WordPress, Shopify, Wix, Webflow, API. Astro is not named. Our marketing site is Astro, static, zero-JS default, dispatch-only (WAYS-OF-WORKING.md / CLAUDE.md #14). Auto-publish does not plug in.
Do not adopt. The site already has seven FSA-checked compliance guides (outputs/provenbatch-site/src/pages/guides/). gtm-channels-research.md §4 already has an eight-item human-written shortlist (Bake Diary alternative, PPDS checker, home-baker pillar, live matrix generator, honest FoodCore comparison, batch-vs-recipe manifesto, "best of" listicle, PAL guide). Thirty machine articles a month would compete with that, violate the site's performance law (≤14 KB critical CSS, Lighthouse ≥95), and — this is the one that is unique to us — put life-safety copy (allergens, Natasha's Law, may-contain) through an unreviewed generator. PRODUCT.md's "restraint is the feature" and "claims are verifiable facts" is a product commitment, not a style preference. Google's scaled content abuse policy (developers.google.com/search/docs/essentials/spam-policies, still current 28 Aug 2026) names "using generative AI tools to generate many pages without adding value" as the example, and since 15 May 2026 those policies explicitly cover AI Overviews / AI Mode as well as blue links.
Backlink exchange (4,000+ "vetted partner sites")
No, and we should not try.
Do not adopt. Google's link-spam policy lists "excessive link exchanges ('Link to me and I'll link to you') or partner pages exclusively for the sake of cross-linking" and "using automated services to create links" as violations. A 4,000-site partner network whose selling point is "$800/mo of backlinks on autopilot, no outreach" is that pattern with a brochure. For a compliance product whose brand is the label proves what went in, a link-scheme penalty (or even the appearance of one) is worse than no ranking.
Reddit / Quora agents
Reddit half is forbidden. Quora half is unused.
Do not adopt. CLAUDE.md #18 / outputs/gtm/social-accounts.md: Facebook, Instagram, TikTok, YouTube — Reddit is dropped permanently. An agent that comments on Reddit to get cited by LLMs is exactly the channel we closed, plus automated engagement, which §19b already records as the line platforms actually police.
GEO audit + "are we cited in ChatGPT?" tracking
Yes, cheaply, by hand.
**Adopt the question, not the product.** Once a quarter, ask ChatGPT / Perplexity / Gemini the five commercial queries we actually care about ("allergen label software UK", "Natasha's Law software", "PPDS labelling software", "FoodCore alternative", "Bake Diary alternative") and write down whether we appear. That is an hour, not $99/month. Schema / answer-first writing for those pages is already in gtm-channels-research.md §4e.
22c. Why the price is the wrong comparison, even if the tactics were clean
$99/mo is 2.5× our top published plan (£39). Paid-ads research in gtm-channels-research.md §5 already closed generic paid as a primary channel because SMB SaaS CAC sits 2–7× above allowable CAC at £9–39. Buying an SEO autopilot at that price, in invite-only beta, to feed a static site that is not on their CMS list, to produce volume we have already decided not to publish, does not pencil. Scale at $299/mo is a specialist we do not need; the agency white-label is a different business.
Their case studies (Myhair.ai, Crono, Influee, Brass & Steel, etc.) are mostly DTC / generic SaaS on WordPress or Shopify. None is a UK regulated-food or life-safety publisher. The numbers are vendor-reported; we did not re-verify them, and we do not need to — the fit fails before the proof.
22d. What is worth doing (already on the books)
Nothing new to file. The in-house version of the useful half of this product is the existing SEO shortlist in gtm-channels-research.md §4f, the seven live guides, directory listings (Capterra / GetApp / Software Advice are free — FoodCore is on them and we are not), and the quarterly "does ChatGPT mention us?" check above. That is the replicable part: few, sourced, FSA-checked pages that a human stands behind, not 30/month on autopilot.
Recommendation: do not buy BabyLoveGrowth, and do not clone its machine. Revisit only if (a) the marketing site moves onto a CMS they actually support and (b) organic search is the primary acquisition channel with a person whose job is to review every published page and (c) the backlink-exchange half can be switched off — (c) is not offered today. Until then the right growth work is the shortlist we already wrote.
23. Third-party review credibility — Trustpilot vs the alternatives (#922, researched 7 Sep 2026)
Dave's instinct is that ProvenBatch will eventually need a third-party review profile (Trustpilot or similar). This is the when-and-how, not a signup. The social-channels precedent applies: Reddit was tried and dropped because it generated upkeep Dave does not have time for (final, CLAUDE.md #18). A review platform that needs soliciting, moderating and answering is the same shape of work.
Trustpilot's own pages 403'd this session (bot wall). Pricing is taken from the live business.trustpilot.com/pricing fetch of 7 Sep 2026; competitor review counts are search extracts plus our own earlier CompliChef pass, flagged where thin.
23a. What the platforms actually cost, and what they demand of Dave
Claim the provenbatch.co.uk profile. Anyone can already write an unclaimed review against the domain.
Reply to every review (a 1-star from someone who wanted a till system is still public). 50 invites/month is enough at our volume. Cannot customise the profile or run widgets on the marketing site without paying.
The only Trustpilot tier that could ever make sense.
Trustpilot — Starter
$99/mo per domain, billed annually up-front, 12-month lock-in (~$1,188). New small-business customers only, ≤ $5M revenue. 100 invites/mo, 2 widgets.
Same claim, plus a prepaid year.
Same replies, plus a widget to keep working on a zero-JS site, plus a contract to cancel.
Do not buy. $99/mo is 2.5× our top plan, annual, for invitation volume we do not have customers to fill.
Trustpilot — Plus / Premium
$319 / $799 per domain per month, annual.
Sales quote.
Dedicated success manager at Premium.
Not a conversation.
Google Business Profile
Free. Reviews, replies, photos, no invitation product.
Needs a verifiable location. Service-area businesses can hide the address. #267 trading address has still never had a mail-forwarding test (outputs/gtm/social-accounts.md).
Reply to Maps reviews. Category choice decides what Gemini thinks we are.
Wrong surface. GBP is a local-pack / Maps tool. We sell national software, not a shop people walk into. A GBP that looks like a bakery is a category lie; a GBP for David Biley trading as ProvenBatch at the trading address is a business listing, not a product-review profile.
Capterra / GetApp / Software Advice
Free basic listing. Paid PPC from ~$2/click with ~$500/mo floors — already rejected in gtm-channels-research.md §4d. G2 reportedly folded Capterra in early 2026 (extract-only; not re-verified).
Create a vendor profile, software category, link the site.
Occasional review responses. No invitation lock-in. Software buyers already look here; FoodCore is on the majors and we are not.
The right first platform, and it is already on the GTM shortlist as a week-one job. Not a Trustpilot substitute — a different buyer.
Paid Trustpilot's contract term is the part the brochure underplays: "Our initial contract term is 12 months, prepaid" (FAQ on the live pricing page). That is a year of $99–$319/mo before we know whether anyone in this market uses Trustpilot to choose labelling software.
23b. Do we have enough customers for a profile to look like proof?
No. Invite-only beta. outputs/gtm/beta-pipeline.md scoreboard, 7 Sep 2026: 6 independent signups (Alison, Minimixers, Bake It Or Leaf It, Little Bakeshed, Canal & Crumb, whisked craves), four more still invited. Paying conversion is Dave's call, not a count we have. PRODUCT.md is deliberate: "No testimonials yet, deliberately — the proof strip uses facts until real ones exist."
A profile with 2 reviews reads as "new and unloved", not "trusted". BrightLocal's 2026 consumer survey (via searchlab.nl compilation) puts the "trustworthy rating" threshold around 40+ reviews; even if that number is local-services-shaped, single digits are the danger zone, not the start of a ladder.
CompliChef is the cautionary neighbour. RESEARCH.md §3 already recorded "~5 Trustpilot reviews" on 19 Aug 2026 for a solo-founder food-safety tool ~1 year in. Search extracts on 7 Sep 2026 still show a sparse profile (one UK extract: TrustScore 3.5 from 1 review; later extracts show a handful more). That is a year of being on Trustpilot and still looking small. Direct Trustpilot fetches were bot-walled this session, so treat the count as extract-only; the shape is not in doubt.
FoodCore's credibility play is directories (Capterra / GetApp / G2 / Software Advice / SaaSworthy — gtm-channels-research.md §4d), not Trustpilot. No foodcore.io Trustpilot profile surfaced in search. AllergenKit likewise: SEO guides and free tools, no review-platform presence we could find. The comparable tools that look grown-up on review sites are hospitality/enterprise (Kafoodle on Capterra, Food Label Maker 24 reviews on G2) — a different buyer.
gtm-channels-research.md §4e: review platforms carry a ~3× citation multiplier in AI answers. That is real, and it is an argument for having a profile once it is dense, not for opening an empty one so ChatGPT can cite "ProvenBatch (2 reviews)".
23c. The upkeep, honestly
The free Trustpilot half is not free in time:
Soliciting. Trustpilot and ASA/CAP both forbid incentives for reviews. The legal way is "if you have a minute, here is the link" after a real result (first successful label print, first EHO visit survived). That is a support-email habit, not a widget. Dave is the support function.
Responding. Every review is public under the company name. A beta tester who hits #1128's unit-cost bug, or anyone who expected a bakery POS, leaves a 1-star that sits there until someone answers it in public, in Dave's voice. CLAUDE.md #12 already exists because a session must not put words in front of a customer under his name; a Trustpilot reply is the same risk with a worse URL.
Moderation. Flagging fakes, chasing Trustpilot when a competitor or a confused visitor reviews the wrong product. Low volume now; not zero.
That is the Reddit lesson with a 12-month invoice attached.
23d. Recommendation
Do nothing on Trustpilot yet. Do not claim a profile, do not pay, do not put a TrustBox on the site, do not build an in-app "leave us a review" prompt.
Do the free Capterra / GetApp / Software Advice listings when the marketing site is the public face of a product someone can actually buy (GA, or the moment /beta flips to a trial). That work is already item 1 of the week-one list in gtm-channels-research.md §4d. It is a software-buyer surface, it costs £0, and FoodCore being there while we are not is a real gap. Not filed from this issue — it already has a home.
Trigger that would change the Trustpilot answer: GA has happened and we have ≥ 20 paying customers (not testers) and Dave will commit to answering every review within a week. Then claim the free profile and ask those 20, once each, by hand. Do not upgrade to Starter until the free profile has enough reviews that a widget would not advertise emptiness, and branded Google results are actually showing the Trustpilot snippet. Revisit earlier only if a specific procurement form asks for a Trustpilot URL — that is a one-customer exception, not a strategy.
Google Business Profile is a separate, weaker question (David Biley trading as ProvenBatch at the trading address, once #267 mail-forwarding is proven). It is not how software gets chosen and it is not a substitute for the above.
No follow-up issue. The trigger is the follow-up.
Sources (fetched 7 Sep 2026): business.trustpilot.com/pricing (live; 12-month prepaid in the FAQ) · gtm-channels-research.md §4d–4e · PRODUCT.md Evidence on Hand · beta-pipeline.md scoreboard 7 Sep 2026 · RESEARCH.md §3 CompliChef · social-accounts.md #267 · BrightLocal 2026 consumer survey via searchlab.nl/en/statistics/google-business-profile-statistics-2026 (the "40+ reviews" figure; local-services-shaped, used only as the danger-zone illustration) · search extracts for complichef.co.uk Trustpilot (direct fetch bot-walled).
Notes
Last updated 7 Sep 2026
24. "How much should I charge?" — discovery gap plus a one-line invert (#919, researched 7 Sep 2026)
The claim in the issue is that "how much would you charge for X" is a very common question in small-food-business groups, and that ProvenBatch already computes the inputs a pricing answer needs but stops at showing margin on a price the user typed. Both halves check out. This is not a group scrape — it is the public residue of those threads: the blog posts, free calculators and competitor guides that exist because the question is asked on a loop.
24a. What is actually being asked, and what people answer with
The question is almost never "what is my food cost %". It is "what do I put on the sticker": a dozen cupcakes, an 8-inch birthday cake, a loaf, a market-stall brownie. The answers groups give each other cluster into four folk rules, all visible in the 2025–2026 pricing-guide literature that exists to argue with them:
Folk answer
What it actually is
Why it fails
"3× ingredients" / the 3-2-1 rule
Old caterer's heuristic: charge three times the ingredient bill. PastryCal still has to explain it in 2025 to say stop using it.
Ignores labour, which on a decorated cake is the cost. Bakingsubs tracked 47 birthday cakes: ingredients ~$14.60, labour 3.2 hours.
Copy the Facebook / Instagram neighbour
Market-rate matching.
The neighbour is often the one asking the same question. FoodCore's UK guide (15 Jun 2026) says this out loud: market rate is a sense-check, not a substitute for cost.
A floor on ingredients only. A £3 ingredient bill becomes a £10 sticker before anyone has been paid.
(Ingredients + labour + overhead) ÷ (1 − target margin)
The formula every calculator now ships. Target margin 30–50% home, 50–70% custom, 15–30% wholesale (MyBakeCalc).
Right shape. The fight is over whether the three inputs are real.
Labour is the argument. "Charge what you're worth" is the motivational version of the same gap: people know ingredients, guess time, invent an hourly rate, skip overhead, then undercut because the resulting number "looks high". Wholesale vs retail (typically 50% of retail) is a second question the same threads spawn.
We did not scrape a named Facebook group. D-17 / gtm-channels-research.md §2 already records that the high-value baker groups are coaches' funnels and that posting into them is the wrong move; the question-shape is evidenced by the industry of pages written to answer it.
24b. How the adjacent tools handle it
They treat this question as the funnel, not a feature buried behind a sale price field.
BakeProfit is built around it. No-signup recipe-cost and cake-pricing calculators (/tools/recipe-cost-calculator) take typed ingredients + labour + overhead, let you set a 50–100% markup, and emit a suggested selling price per serving. Free plan $0; Pro $6.99/mo. Their blog library is the same question in 44 product-shaped pages. gtm-channels-research.md §4c already flagged BakeProfit's no-signup calculators as "their whole funnel".
FoodCore published a UK-native version on 15 Jun 2026: foodcore.io/blog/how-much-should-i-charge-for-a-homemade-cake. Four components, the ÷ (1 − margin) invert, worked £ examples (£80 for a 6-inch celebration cake once labour at £15/hr is in). They are sitting on the exact SERP we would want. They do not, from our existing product analysis, invert target-margin → price inside the app the way Recipe Cost Calculator does — the guide is the play.
Recipe Cost Calculator (recipecostcalculator.net, since 2012) is the honest invert: "enter a sell price to see the margin, or a target margin to see the required sell price", plus wholesale / retail / distributor scenarios side by side.
ReciPal (already in insights-bi-research.md §3.1): cost-over-time with target margins.
Craftybase (insights-bi-research.md §3.2): "pricing guidance auto-updating when material prices change" — the live-cost half of what we already do, plus a suggested price on top. $49–349/mo, a different buyer.
The pattern: the calculator that suggests a price is the marketing site; the app that costs from real packs is the product. BakeProfit and FoodCore have the first. We have the second.
24c. What we already do, honestly
Verified against recipe-product-variance-design.md §2.1a, recipe-product-cost-ownership-spike-2026-09.md, and USER-GUIDE.md (Products: sale price, margin as sale price minus linked recipe unit cost).
Cost is real, and that is the moat.lineCost() → recipeComputed().cost → productUnitCost() (recipe unit + packaging unit since #1256), from the supplier pack actually on the ingredient, not a typed "I think butter is £1.80". Live. Recomputed every render.
Margin is one-way. The user types products.unit_price. The drawer and the Products list show margin £ and %. There is no target-margin field and no suggested price. The last mile of every calculator above is missing.
Labour is a time, not a cost. Recipes carry Batch time (one end-to-end figure in minutes). There is no hourly rate, so we cannot honestly turn that into a labour line. Inventing £15/hr to match FoodCore's blog would be a new product pretending to be a feature.
Overhead is the other ledger. #340: planned unit economics and actual cash (expenses, drawings, mileage, home-working, other_income) deliberately do not reconcile. Folding a made-up overhead % into a suggested sticker would mix the two systems the design forbids mixing.
The existing guide is the cost half, not the price half.outputs/provenbatch-site/src/pages/guides/costing-a-recipe-moving-prices.md (7 Aug 2026) explains why last month's butter price eats the margin. It never answers "so what do I charge".
So: capability gap on the last mile (one-way margin), discovery gap on the question people actually type into Google and Facebook, and a hard stop on labour/overhead because we do not have honest inputs for them. The gap is not "the app already does this once they find it".
24d. Recommendation: content now, invert later, no pricing consultant
Both, scoped. Neither is this issue.
Content play (marketing site, docs-only, Later). A UK-native guide that occupies the SERP FoodCore took in June: price from the batch, not from Instagram. Use the folk-rule table above, show why 3× ingredients fails, and point at the Products margin on a real pack cost. Extend costing-a-recipe-moving-prices.md or add a sibling under /guides/. Do not ship a no-signup calculator that invents ingredient costs — that is BakeProfit's game and it throws away the moat. A calculator that needs the catalogue is the app. Filed as a follow-up from this issue.
One-line invert on the Product drawer (app, Later, effort: low). Next to sale price: "at n% margin this wants £X.XX" — unit cost / (1 − n/100), using the existing productUnitCost() figure (recipe + packaging). Optional target % with a sensible default (e.g. 50%), or just show 40 / 50 / 60 as quiet chips. Tapping a chip can fill unit_price; it must not silently overwrite a price the user already typed. Out of scope for that issue: labour, overhead, hourly rate, wholesale/retail split, "complexity multipliers", any new column. Those wait on an hourly-rate decision we have not made. Filed as a follow-up.
Do not build a pricing engine. No "3× ingredients" mode, no FoodCore-blog replica inside the app, no overhead slider. The folk rules are what we are replacing, and the two ledgers stay unreconciled.
25. Security certifications — Cyber Essentials is the first rung, not now (#832, researched 7 Sep 2026)
Full plan with triggers: outputs/docs/security-certification-plan.md. Positions below are proposed; Dave confirms. Nothing has been bought or claimed.
Proposed
Trigger
Cyber Essentials
First rung. Official IASME micro fee £320 + VAT / year. Not now.
A named customer / supplier form asks, or GA + ≥ 20 paying customers.
Cyber Essentials Plus
Later. ~£1,350–£1,500 + VAT for 1–9 users.
A named customer or tender requires Plus, in writing.
ISO 27001
Not now. All-in year one for a 1–10 person UK SaaS commonly £8–25k.
A deal whose first-year value is ≥ 3× that cost, or a named procurement that will not take CE.
SOC 2 Type II
Not now, and US-shaped. $25–50k year one. D-01 is UK only.
We reverse D-01, or a specific US/enterprise deal names it.
#211 is closed and already left the trust page to this issue. The cheap half (SECURITY.md, /.well-known/security.txt, a /security page listing only controls that exist) is specified in the plan and not built here — a follow-up, after Dave confirms the table. Do not claim tamper-evident safety records on that page until audit F-01/F-02 are fixed.
The DPA already tells the truth: we do not hold ISO 27001 or SOC 2. Keep saying so.
Sources: plan §5.
Notes
Last updated 16 Sep 2026
30. Backlinks for provenbatch.co.uk — earn them, never buy them, and no tool is needed (researched 16 Sep 2026)
Dave asked whether the marketing site needs a backlinking effort, prompted by the steady stream of apps offering "backlinks → traffic → customers", and whether a dedicated tool is warranted. This section is the answer on the record. It extends §22 (BabyLoveGrowth, 7 Sep 2026), which scored one such product and rejected its backlink exchange; nothing here reopens that verdict.
30a. What a backlink actually gives us
A link from another site to provenbatch.co.uk does three distinct things, and they are worth separating because the vendors blur them:
A ranking signal. Google treats a link from a relevant, established site as a vote that the domain is real. For a domain first published in Aug 2026 this matters more than for an old one: gtm-channels-research.md §4f puts a new domain at 3–6 months of suppression, and relevant links are one of the few things that shorten it.
Referral traffic in its own right. A link on a page people already read (a council's home-baker page, a group's pinned resources) sends visitors whether or not Google ever ranks us.
AI-answer citations.gtm-channels-research.md §4e: review platforms and directories carry a ~3× citation multiplier in ChatGPT / Perplexity / Gemini recommendation answers, and third-party comparison content is cited ~3× more than us-vs-them pages.
All three compound slowly. The expectation in §4f still stands: near-zero Google traffic before month four regardless of links. Set against the live evidence (outputs/gtm/beta-pipeline.md §C: six signups to 14 Sep, at least four credited to the two Cake Shed Facebook groups, none to search), backlinks are a slow-burn investment for 2027, not a lever for the 1 Oct launch.
30b. Why the apps are the wrong way to get them
Google's spam policies (developers.google.com/search/docs/essentials/spam-policies, page dated 28 Aug 2026, re-read 16 Sep 2026) list as link spam, quoted: "Exchanging money for links, or posts that contain links", "Excessive link exchanges ('Link to me and I'll link to you') or partner pages exclusively for the sake of cross-linking", "Using automated programs or services to create links to your site", and "Low-quality directory or bookmark site links". That list is the product description of nearly every backlinking app: a partner network, a marketplace of paid placements, or an automated outreach engine. Two consequences specific to us:
The penalty is asymmetric. A product whose brand is the label proves what went in cannot afford even the appearance of a link scheme (§22b's reasoning, unchanged).
The CMS assumption fails anyway. The tools assume WordPress or Shopify with auto-publish; the site is static Astro, dispatch-only (D-23, AGENTS.md #14). §22a's CMS finding applies to the whole category, not only BabyLoveGrowth.
Do not buy a backlinking tool or service, at any price. Revisit only under §22d's three conditions, and note that (c) — the ability to switch the exchange off — is the one no vendor offers.
30c. Where the site stands, and the three gaps
The technical foundation is in place (outputs/provenbatch-site/, verified 16 Sep 2026): robots.txt open and advertising /sitemap-index.xml; the sitemap generated by @astrojs/sitemap with /beta/welcome excluded by design (#733); canonical, hand-set meta descriptions and a full OpenGraph set in src/layouts/Base.astro; eight FSA-sourced guides (src/data/guides.js); six sector pages under /for; a UK-competitor /compare (#1911); Cloudflare Web Analytics live since 15 Sep (#1873, cookieless per D-27). Three things are missing, and they are the whole job:
Gap
Evidence
Why it matters for links
No Google Search Console property
grep -ri "search console" across the repo: 0 hits; no google-site-verification token in src/ or public/. ⚠️ Amended later the same day: public DNS does carry a google-site-verification TXT at the apex (and stellaapps.com has its own) — almost certainly Google Workspace's domain verification, since both domains have Google MX and the support@/social@ aliases. It is not evidence of a Search Console property, but it may let the Domain property verify with no DNS change; see outputs/gtm/drafts/search-console-dns-and-register-1915-2026-09-16.md §0
It is the only free tool that shows who links to us and what we rank for. Without it every backlink question is unanswerable, and any tool a vendor sells is reselling this data
The free directory listings were never made
Not in HANDOVER, the beta pipeline or any issue; gtm-channels-research.md §4d called them "a week-one job at site launch"; §23 tied them to "the moment /beta flips to a trial" — /trial shipped under #1818
These are the only legitimate backlinks available on demand, and the ones with the AI-citation multiplier. FoodCore is on all of them; we are on none
No structured data
Zero ld+json, microdata or @context anywhere in the site (the repo's only JSON-LD hits are the recipe importer consuming other sites' markup)Corrected 16 Sep 2026 (#1915):Organization on every page via Base.astro; SoftwareApplication (Starter £9 / Standard £19 / Pro £39) on /; Article with citation on each guide from that guide's existing sources: frontmatter; BreadcrumbList on guides and /for/*. FAQPage is not emitted — none of the eight guides is a discrete Q&A. Helper: outputs/provenbatch-site/src/lib/jsonld.js. Live after a deploy-site.yml dispatch (site: provenbatch).
§4e asked for schema so the guides get cited; the schema half is in the tree.
Directory correction, 16 Sep 2026.gtm-channels-research.md §4d recorded "G2 acquired Capterra" as extract-only, unverified. Capterra's own vendor page (capterra.com/vendors/, fetched 16 Sep 2026) now reads "Capterra, powered by G2 Digital Markets" and its Create a product listing button goes to g2.com/products/new; g2digitalmarkets.com shows the Capterra, GetApp and Software Advice marks together. So the July finding is confirmed and the job is smaller than the research assumed: one G2 product profile is the route in for all four sites. Whether the Capterra/GetApp/Software Advice listings appear automatically from the G2 profile or need a separate step inside the same account was not visible from the public pages — establish it inside the account, do not assume. Serchen and SourceForge remain optional: Google's list names "low-quality directory" links as spam, and neither carries the review-platform citation multiplier, so they are worth a listing only if it is free and takes minutes.
30d. How to go about it — in order, and what each costs
Verify provenbatch.co.uk in Search Console as a Domain property. Google's own guide (support.google.com/webmasters/answer/9008080, read 16 Sep 2026): a Domain property covers every protocol and subdomain, verified by one DNS TXT record on the apex (@) whose value Search Console generates in the form google-site-verification=…. The record goes in the Cloudflare DNS zone. The sitemap is then submitted once in the Sitemaps report (the robots.txt line already lets Google discover it, but the report is what shows crawl errors). Cost: nil; one sitting. Dave's part: the Google login and the TXT paste. Record the property in outputs/gtm/social-accounts.md alongside the other Google assets. This is item one because every later decision reads its Links report.
Create the G2 product profile, then confirm Capterra, GetApp and Software Advice from inside it. A session drafts the full listing (category, one-line description, long description, feature checklist, pricing tiers from pricing.astro, screenshots from the existing site assets) so each field is a paste. Cost: nil for the basic listing (paid PPC from ~$2/click with ~$500/mo floors — already rejected in §4d, still rejected). Dave's part: the account and the pastes. No review-collection push yet: §23's Trustpilot reasoning applies here too — an empty profile is a listing, a dense one is a credential, and the second comes after GA.
Add JSON-LD to the site.Organization + SoftwareApplication on the home page, Article (with citation for the FSA/GOV.UK sources the guides already list) on each guide, FAQPage only where a guide genuinely answers discrete questions, BreadcrumbList on guides and sector pages. Emitted from Base.astro / GuideLayout.astro from the same props the meta tags already use, so it cannot rot separately (D-23's reason for Astro). Not a backlink; it is the half of §4e that makes the existing guides citable. Cost: one session, no Dave step.Shipped 16 Sep 2026 (#1915): the helper and the layout emission are in the tree; see the corrected §30c row. Dave still owns items 1–2.
Build the two link-earning free tools already on the shortlist (gtm-channels-research.md §4f items 2 and 4): the interactive "Is my food PPDS?" checker and the live allergen-matrix generator. Councils and hobby blogs link to things that help their businesses comply; they do not link to pricing pages. The council-signposting precedent is already on record (§12: West Norfolk's cake-makers PDF; start with Buckinghamshire). With the "Bake Diary alternative" page removed by Dave on 31 Jul 2026 (provenbatch-website-content-foundation.md §6e), these two carry the whole earned-link load. Cost: two build sessions each. Not filed here — they are product work with their own specs.
Ask for a link only where a relationship already exists, one honest request each and never at volume: the Cake Shed group admins (D-17's admin-partnership route), NMTF, the Cake International exhibitor listing, a supplier whose receipt format we already parse. Drafts go through the existing outreach-upkeep routine's drafts-only, capped, Dave-sends rails.
Fold "who links to us?" into the quarterly check §22b already prescribes ("does ChatGPT mention us?"): one hour, Search Console's Links report plus the five commercial queries.
Items 1–3 are filed as #1915. Items 4–6 are the existing plan; nothing new to file.
30e. What is deliberately not proposed
No Google Business Profile — §23 rejected it as the wrong surface for national software; the Metricool research's open question stays open, not re-raised.
No Trustpilot — §23's trigger (GA, ≥20 paying customers, a weekly review-answering commitment) has not been met.
No guest-post placements, no "DA/DR" score services, no link marketplaces. Every one of them is on Google's list above, and a domain-authority number is a vendor's proxy metric, not anything Google publishes.
No fifth social channel (AGENTS.md #18).
Sources (16 Sep 2026): developers.google.com/search/docs/essentials/spam-policies (page dated 28 Aug 2026) · support.google.com/webmasters/answer/9008080 (Domain property, DNS TXT) · developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap · capterra.com/vendors/ · sell.g2.com and g2.com/products/new · g2digitalmarkets.com · internal: §4e/§4f/§4d of gtm-channels-research.md, §22, §23, outputs/gtm/beta-pipeline.md §C, outputs/provenbatch-site/ as inspected, go-to-market-decisions.md D-23/D-27.
Notes
Last updated 20 Sep 2026
32. Pepesto — a supermarket price API, not a food-business tool (researched 20 Sep 2026)
Dave's brief: pepesto.com "looks like something that could be hugely valuable" — research it in detail and find every place it could add real value to our users. Full write-up, graded and sourced: outputs/docs/pepesto-integration-research-2026-09.md. Decided the same day: recorded here and in the admin research library; no issues filed yet — §6 of the doc is a candidate list, not a backlog.
What it is. Pepesto Solutions GmbH (Adliswil, Switzerland, founded 2023, two founders, no disclosed funding) sells a live supermarket price and catalogue API — 28 chains in 13 countries, daily refresh with promotions — plus a consumer recipe-to-basket engine and a self-hosted MCP server. UK: Tesco, Sainsbury's, Asda, Morrisons, Waitrose, Ocado, about 4,000 cooking-ingredient SKUs a chain. Not Aldi, Lidl, Costco or Booker. Ireland: Tesco IE, Dunnes, SuperValu. Pay-as-you-go €29.90 credit packs, or €79/month on a six-month commitment for half-price calls; re-pricing 50 known product URLs costs €0.96, a shopping-list match €0.32.
The finding that corrects the premise: its product data carries no ingredient list, no allergen declaration, no per-100g nutrition and no barcode — verified against the CatalogProduct / IndexedProduct schemas in their openapi.yaml, not inferred from copy. The homepage's "nutrition data" is four recipe-level integers from their knowledge graph. So it gives the compliance core nothing, and decision 432 (no external lookup for label content) and #28 (no external nutrition lookup without a design record) are untouched.
Where it does add value — the money half, ranked: (A) a supermarket-bought pack cost kept true before the receipt, through the same price-drift chip priceDrift() already shows — and the missing caller for the deployed, never-wired check-prices function; (B) the saved shopping list priced at today's shelf price with offers, optionally "cheaper at Asda this week" — nothing in §3 has live supermarket pricing; (C) onboarding prefill of pack size, price and URL for a chain-bought supplier product, with declaration and allergens still from the pack photo. Later or experiment-only: a promotions watch on Today, Ireland for free, and a free list-to-basket hand-off that checks out inside Pepesto's own consumer app.
Not a fit, recorded: nutrition (#28 — wrong shape of data; CoFID would be the candidate if a lookup were ever chosen), declarations and allergens (no data, and 432), recipe import (parse-recipe + fetch-page already do URL/photo/PDF/text), shelf analytics (for brands listed in supermarkets), the MCP and agent-to-cart (consumer products).
Three constraints on any yes: their API terms allow three days of retained price history without a Partner agreement (one observation per sourcing is fine, a trend line is not); a new sub-processor row and a credential-inventory entry before the first Deno.env.get(); and the Bake Diary posture — a three-year-old two-founder vendor, so every surface degrades to today's behaviour with the API absent and no stored figure depends on it. Pilot on credits, per tap, metered through api_usage_log; do not sign the six-month plan first.
Notes
Last updated 22 Sep 2026
33. Distribution channels that need no content treadmill — validated and ranked (researched 22 Sep 2026)
Dave's brief: which distribution channels exist for the app that do not require content creation but could build customer numbers, ranked by impact. Full report, graded and sourced: outputs/docs/distribution-channels-research-2026-09.md. It filters §12/§30/§31 to the no-content half and validates the three lanes nobody had checked — app stores, integration marketplaces, free software directories. A channel qualifies if it keeps working with nobody writing posts, filming video or producing articles; a one-off listing, a build, an email or an ad headline is in scope, a guest masterclass or a social queue is not. That definition is what reorders the 90-day plan: three of §C11's seven priorities are content in disguise.
The app already qualifies for Google Play today.manifest.webmanifest (standalone, id/scope/start_url, 192/512/maskable, a shortcut) + sw.js + an HTTPS host is exactly what a Trusted Web Activity needs, alongside Digital Asset Links proving app and site share a developer (developer.chrome.com, verified). Registration is US$25 once, no renewal (Play Console help, verified). Because a TWA renders the live site, app updates never go through Play review — only the shell does. Personal accounts created after 13 Nov 2023 must pass a testing requirement first (verified; reported as 12 testers for 14 days, with organisation accounts exempt — establish that inside the console, do not assume).
🔴 The landmine is the plan chooser, not the wrapper. Play's Payments policy lists "cloud software and services" among purchases that must use Play billing, and exempts only goods that "can only be consumed outside of a Play-distributed app". Our app is the consumption surface, so the Android build must carry no purchase path — trial signup yes, PRICES_GBP plan cards and the Stripe hand-off suppressed, subscribing on the web. Design that in; do not discover it at review. (US alternative billing starts charging fees 1 Oct 2026 — US-only; the CMA opened its steering conduct-requirement consultation 30 Jun 2026, so UK linking rules are moving.)
The Xero App Store is closed to us on price. Revenue share retired; from 2 Mar 2026 the tiers are Starter free (5 connections), Core $35 AUD (50), Plus $245 AUD/mo — the cheapest tier where a listing is even optional — Advanced $1,445 AUD, Enterprise by application (listing mandatory), all with certification and $2.40 AUD/GB egress overage (developer.xero.com/pricing, verified). A listing costs more per month than three Pro subscriptions. QuickBooks (three reviews
annual re-review), Zapier (needs a public API we do not have) and Shopify/Square (build a platform app for a minority of our buyers) all fail too: integration marketplaces charge in engineering, and two now also charge rent.
Microsoft Store now costs nothing — "no registration fees for either account type" via storedeveloper.microsoft.com (verified). Wrong device for a market trader, but it is a free ride on the same PWA package, so 20 minutes after Play.
Apple stays deferred: $99/yr, individual enrolment publishes the enrolee's personal legal name as the App Store seller (collides with D-07's published-vs-verified rule), organisation enrolment needs a D-U-N-S number, and guideline 4.2 rejects wrappers with no native capability — for a surface iOS users can already install from Safari. Trigger: Play shows installs and a D-U-N-S exists, or the CMA's iOS interoperability commitments make the wrapper unnecessary.
The ranking (impact first, effort priced separately): 1 Play as a TWA (+ Microsoft Store); 2 member-benefit and perk listings plus a partner affiliate link — NMTF /deals still carries no software; 3 the two-sided referral credit and #688's public QR allergen menu, the only compounding row; 4 #1944 G2 → Capterra/GetApp/Software Advice, free and already packaged; 5 Search Console plus Bing Webmaster + IndexNow (free, never on the list, and Bing feeds several AI assistants); 6 exact-match Google ads, because they price a trial inside 30 days; 7 free long-tail directories — AlternativeTo as the Bake Diary alternative, SaaSHub, one Product Hunt afternoon; 8 the Erudus partner listing (nil in 90 days, high in 2027); 9 Brother/MUNBYN/OLBAA partner programmes. No row replaces activation and word of mouth — rows 4–7 are hours of hygiene, rows 8–9 are emails whose payoff is 2027.
Rejected and recorded so they are not re-proposed: the four integration marketplaces, Apple for now, Trustpilot (§23's trigger unmet), Google Business Profile, AppSumo and lifetime-deal marketplaces (a lifetime price at £9–39/mo economics destroys the subscription it is meant to seed, and the buyers are worldwide software collectors, not UK PPDS businesses), paid directory PPC, backlink marketplaces, reseller/white-label and hardware lock-in, a fifth social channel.
Not filed: Play + Microsoft Store distribution, the referral credit, Bing/IndexNow, and shareable EHO-pack/matrix links are candidates, not issues. #1944 and #688 already exist.
Notes
Last updated 3 Oct 2026
36. Farm shops — the rules the app does not yet cover, and the till and scale landscape (researched 3 Oct 2026)
The finding: a farm shop is served today for what its own kitchen makes, but its counters, own produce and bought-in range bring in rules the app has never modelled. None of the rules below was recorded here before. The full gap analysis, ranked for a pilot with named shops, is outputs/docs/farm-shop-fit-review-2026-10.md; filed as #3055 (children #3056–#3072).
Beef labelling is compulsory for everyone selling fresh or frozen beef or veal, farm shops and butchers included. A label (except on mince and trimmings) needs a reference number or code, and born in / reared in / "Slaughtered in:" / "Cutting/cut in:". The slaughterhouse and cutting-plant licence numbers do not apply to beef sold loose over the counter. Every business selling beef must keep a traceability system whose records link the reference numbers on its labels to its beef intake. Claims such as farm name are voluntary; breed, maturation and "grass fed" need approval. (GOV.UK, compulsory beef labelling scheme; GOV.UK, set up a traceability system for beef and veal.)
Eggs sold direct from the farm need not be graded or stamped when the conditions are met. The thresholds are 50 hens for selling ungraded eggs at a local public market, and 350 hens for registering as a production site. The pack still carries a best before date (at most 28 days from lay) and "keep refrigerated after purchase". Depending on the route, it also needs the production method, producer code and country of origin. Selling through a shop rather than at the farm gate changes which conditions apply, so check before printing an egg label. (Business Companion, egg producers selling directly to consumers; Bromley and Surrey Heath council guidance.)
Unit pricing (Price Marking Order 2004) applies to goods sold loose from bulk (fruit and veg, meat, fish) or required to be marked with a quantity. Shops with a retail area under 280 m² are exempt from showing the unit price, but must still show the selling price. Most farm shops are under that size, so a £/kg label is mainly an operational convenience rather than a legal duty. (GOV.UK Price Marking Order 2004 guidance; Business Companion, providing price information.)
Butchers' and farm shops' own prepacked lines are PPDS (Natasha's Law); loose counter meat is not, though its allergen information must still be available. This confirms §4 and the butcher note in outputs/docs/gtm-acquisition-research-2026-09.md. (FSA, PPDS allergen labelling changes for butchers.)
The till and scale landscape. Farm shop tills are food-specialist EPOS systems (ICRTouch TouchPoint, Cunninghams Quantum running on Avery Berkel scales, CSY) or general ones (Square, Shopify POS). Counter scales either price the item straight onto the till or print a price-embedded barcode label to scan. General tills handle price-embedded barcodes poorly: Shopify POS needs a third-party app that copes badly with fractional weights. The specialists already sell "Natasha's Law labels" from the scale's label designer, which is the overlap ProvenBatch has to answer with a PLU and ingredients export rather than a rival scale. (ICRTouch farm shops page and brochure; Speciality Food, a guide to point of sale systems; Square community thread on price-embedded barcodes.)
Local supply is the farm shop's identity, and its due-diligence load. Runcton Farm Shop lists 70+ suppliers within 15 miles, and Chatsworth 30+ UK suppliers plus smaller producers. Those suppliers are mostly the jam makers and bakers ProvenBatch already serves, which makes the shop a channel as well as a customer. (FarmingUK directory; Chatsworth farm shop pages.)
Sources: links in outputs/docs/farm-shop-fit-review-2026-10.md § Sources.
Notes
Last updated 3 Oct 2026
37. Selling another maker's food — who answers for its label, and what the shop must keep (researched 3 Oct 2026)
The finding: the maker is responsible for its own food information, but the shop must not sell food it knows or presumes is non-compliant, and its legal defence is reliance on the maker's information plus reasonable checks. So the shop's real need is a dated record of what the maker declared and when the shop checked it. This is the basis for the local-producer declaration request (#3061; design record outputs/docs/producer-declaration-request-design.md).
Retained Reg. (EU) 1169/2011, Art. 8. The operator responsible for food information is the one under whose name the food is marketed (8(1)), and it must ensure the information is present and accurate (8(2)). An operator that does not affect the food information "shall not supply food which they know or presume, on the basis of the information in their possession as professionals, to be non-compliant" (8(3)). (legislation.gov.uk, Reg. 1169/2011 Art. 8.)
Food Safety Act 1990, s.21. Due diligence is a defence (21(1)). A seller who neither prepared nor imported the food meets it by proving reliance on "information supplied by" a person not under its control, plus reasonable checks or reasonable reliance on the supplier's checks, with no reason to suspect a problem (21(2)–(4)). (legislation.gov.uk, Food Safety Act 1990 s.21, revised.)Corrected in place, 3 Oct 2026: on its own wording, s.21(2)–(4) covers only offences under s.14 or s.15 (stand-in security review on #3115). For the allergen and labelling offences, the Food Information Regulations 2014 (England) apply it, Sch. 4 Pt 5. That schedule substitutes "regulation 10 of the Food Information Regulations 2014" for "section 14 or 15 above" in s.21(2), so the reliance route is also available as a defence to a reg 10 prosecution, subject to the court finding the checks reasonable. (Reg 12(1) names ss.10, 32, 37 and 39; reg 12(5) applies the Schedule, which adds s.21.) Wales, Scotland and NI have their own FIR instruments, not checked here. (legislation.gov.uk, SI 2014/1855 Sch. 4 Pt 5, revised.)
FSA technical guidance (gov.uk; published 23 Aug 2023, updated March 2025), under Responsibilities (Article 8): as best practice (para 75, not a duty), businesses "should review ingredient information for foods provided by them and ensure that their suppliers provide them with the necessary information". The mandatory para 76 restates Art. 8(8): a supplier of food "not intended for the final consumer or to mass caterers" must give the next operator "sufficient information to enable them … to meet their obligations under paragraph 2". It is silent on how that information is recorded.
PECR still applies to any email we would send a maker. This confirms §19d: sole traders and non-limited partnerships are individual subscribers (ICO, Business-to-business marketing). So a declaration request must come from the shop, or be purely transactional with no ProvenBatch promotion in it.
Anthropic API retention is still deletion of inputs and outputs within 30 days by default (Anthropic Privacy Center, "How long do you store personal data", updated 1 July 2026). This is relevant whenever a third party's photo is read.
Distribution channels that need no content treadmill — validated and ranked by impact
Correction, 29 Sep 2026 (#2666, #2718): the public launch planned for 1 Oct 2026 is postponed with no new date; sign-up stays invite-only (decision #2666). Wherever this document says "1 Oct", "from 1 Oct" or "90 days from 1 Oct" as the opening, read "the public opening (no date yet)". The findings below are as researched on 22 Sep and are not rewritten; the three action items and the G2 gate are reworded in place.
Researched 22 Sep 2026 against live primary sources, three weeks after gtm-acquisition-research-2026-09.md (20 Sep) and six days after RESEARCH.md §30 (directories and backlinks). Dave's brief: "validate what distribution channels exist for our app that don't require content creation but could help build customer numbers, and rank them by impact." This document does not repeat the channel map in gtm-channels-research.md or the 90-day plan in gtm-acquisition-research-2026-09.md §C11 — it takes both as read, filters them to the no-content half, validates the ones nobody had checked (app stores, integration marketplaces, free software directories), and puts the whole set in one ranked order.
How to read the grades (same scheme as the 20 Sep pass). [verified] = the primary page was fetched and read on 22 Sep 2026. [reported] = a secondary source or search extract not opened. [inference] = our reasoning on the evidence. [repo] = a fact from this repository or the production database.
What "no content creation" is taken to mean. A channel qualifies if, once it is set up, it keeps working without anyone writing posts, filming video, or producing articles. A one-off listing (a form, screenshots we already own), a build, an email, or an ad headline is in scope. A guest masterclass, a podcast appearance, a blog estate or a social queue is out — which is the distinction that reorders the existing 90-day plan, because three of its seven priorities are content in disguise (§4b).
0. Headlines — what this pass establishes
The app already qualifies for Google Play today, technically.manifest.webmanifest has id/scope/start_url, display: standalone, 192/512/maskable icons and a shortcut; sw.js is a real service worker; the host is HTTPS [repo]. Trusted Web Activity needs exactly that plus Digital Asset Links to prove the app and the site share a developer [verified]. So the Android listing is a packaging job, not a rewrite — and because a TWA renders the live site, app updates do not go through Play review; only the shell does [inference on the TWA model]. Registration is US$25 once, no renewal [verified].
The one policy landmine is the plan chooser, not the wrapper. Play's Payments policy lists "cloud software and services" among the things that must use Play billing when they are sold in-app, and exempts "purchases of digital goods or services that can only be consumed outside of a Play-distributed app and cannot be accessed in a Play-distributed app" [verified]. Our app is the consumption surface, so the safe posture is an Android build with no purchase path inside it — trial signup fine, PRICES_GBP plan cards and the Stripe hand-off suppressed, subscribing done on the web [inference]. That is a build constraint to design in from the start, not a thing to discover at review.
The Xero App Store is closed to us on price, as of this year. Xero retired revenue share and from 2 Mar 2026 prices developers by tier: Starter free (5 connections), Core $35 AUD (50), Plus $245 AUD/month — the cheapest tier where an App Store listing is even optional — Advanced $1,445 AUD, Enterprise by application (listing mandatory) [verified]. £130-ish a month for a listing, before the integration is built, against a £39 top plan. Dead on arrival, and the 3-active-connections review gate [reported] would bite first anyway.
Microsoft Store now costs nothing to join — "no registration fees for either account type" via storedeveloper.microsoft.com, individual or company [verified]. It is the wrong device for a phone-first food business, but it is a free ride on the same PWA package as Play, so it costs minutes once Play is done.
The highest-value no-content channel is one we have already filed twice and not shipped: the referral credit (gtm-acquisition-research-2026-09.md §B9, unfiled) and #688 the public QR allergen menu (open, Later, effort: high) [repo]. Every other channel here is rented attention; those two are the only ones that compound with each customer won, and neither needs a word of marketing copy after it ships.
Apple stays deferred, and for a reason worth writing down: individual enrolment puts the enrolee's personal legal name on the App Store as the seller [verified], which collides with D-07's published-vs-verified address rule; the organisation route needs a D-U-N-S number, a work-domain email and a public site [verified]; and guideline 4.2 rejects wrappers with no native capability [reported]. $99/year plus a review fight for a surface iOS users can already install from Safari.
1. The ranking
Impact = expected new paying customers over the 90 days from 1 Oct, at ProvenBatch's actual scale, divided by nothing (it is an absolute judgement, not a ratio) — effort and cost are separate columns so a cheap small channel does not outrank an expensive large one by arithmetic. Every impact figure is [inference]; the facts each rests on are graded in §2.
#
Channel
Ceiling on reach
Our cost
Effort to open
Ongoing content
90-day impact
1
Google Play listing as a TWA (+ Microsoft Store on the same package)
Every Android owner searching "allergen label", "food labels", "recipe costing" — the surface where consumer-shaped food-business apps actually get volume (Bakesy: 1,074 Play ratings, 87 UK App Store ratings) [repo]
US$25 once
One build session + Dave's console account, ID check, and possibly a 12-tester closed test
None — the listing serves the live site
Medium-high, and the only one that grows on its own: a permanent, searchable shelf next to zero UK competitors (FoodCore, AllergenKit and CompliChef Pro are web-only) [repo]
2
Member-benefit and perk listings, plus a partner affiliate link — NMTF /deals, Encore Kitchens and Mission Kitchen perks, Farm Retail Association supplier membership, accountants' and coaches' "tools we recommend" lists
Thousands of exactly-right businesses per listing; NMTF /deals carries a bank, the AA, solicitors, an insurer and an accountant and no software [repo]
£0–£150/yr per body (FRA/NCB memberships priced on application)
One email or call each; a promo code
None — the partner's audience, the partner's words
Medium-high but lumpy: one live listing in a trader-intake email beats a month of anything else; nothing may land inside 90 days
3
Referral credit + the shareable allergen menu / EHO pack (#688)
Every customer becomes a surface; the realistic unit is a trader showing the matrix to the trader beside them [repo]
£0 (no referral tooling — the FreeAgent stacking-credit pattern needs none) [repo]
One build session for the credit; #688 is effort: high (public token, edge function, migration)
None, forever
Low now, highest later: with ~6 beta accounts it is arithmetic on a small base, and it is the only channel whose output rises with the customer count
Low-medium, and it buys certainty: the only channel that tells you the cost per trial inside 30 days
7
Free long-tail directories — AlternativeTo (as the Bake Diary and Craftybase/Stocksmith alternative), SaaSHub, Serchen/SourceForge if free, one Product Hunt launch
Intent-matched but thin; a Product Hunt run outside the top 10 is under 500 visitors [reported]
£0 (optional $5 for a faster AlternativeTo review) [reported]
20 minutes each; AlternativeTo needs an account 7 days old [reported]
None
Low: permanent backlinks and AI-citable rows; not a customer engine
8
Erudus integration-partner listing
82,689 caterers and retailers, 211 wholesalers; Planglow and Kafoodle are already listed [repo]
£0 to enquire
One email now; the integration is a 2027 build
None
Nil in 90 days, high in 2027 — it also attacks the onboarding barrier, which is why it is in the plan at all
9
Printer, packaging and label-stock partner programmes — Brother UK's "expert food labelling partners" route, MUNBYN UK dealer/affiliate, OLBAA box inserts
Their customer lists; no UK precedent of a label seller running software inserts [repo]
£0–£50
One email each
None
Low and unproven: Brother's programme is multi-outlet shaped [repo]; worth the emails, not the hope
The honest summary of that table: no channel on it replaces activation and word of mouth. Rows 1–3 are the ones that change the shape of the business; rows 4–7 are cheap hygiene that should all be done in October because each is measured in hours; rows 8–9 are emails whose payoff is 2027.
2. Channel by channel — the validation
2.1 Google Play, as a Trusted Web Activity — open to us, and cheaper than assumed
What is required, verified. Play Console registration is "a US$25 one-time registration fee", with personal and organisation account types; identity verification is mandatory; developers on personal accounts created after 13 November 2023 "must meet specific testing requirements before they can make their app available on Google Play", and new personal-account developers must verify access to an Android device through the Play Console mobile app [verified]. Secondary sources put the testing requirement at 12 testers for 14 continuous days, and state that organisation accounts are exempt from it [reported] — that difference is worth three weeks of calendar, so establish it inside the console before choosing the account type; do not assume it (the same rule #1944 applies to the Capterra question).
On the packaging side, Chrome's own documentation says a TWA's site "and the app are expected to come from the same developer… verified using Digital Asset Links", that TWAs "need to meet the same Add to Home Screen requirements", and names Bubblewrap (the Node CLI) as the tool [verified]; a Lighthouse floor of 80 is widely quoted but was not in the primary page [reported]. Against that, the app is already a real PWA [repo] — manifest, service worker, maskable icon, HTTPS.
The blocker is policy 4.3 / minimum functionality, not the code. Raw WebView shells that mirror a website are the number-one rejection cause; a TWA over a genuine PWA is Google's own recommended shape [reported]. ProvenBatch is on the right side of that line by construction — offline capable, camera capture, install-grade manifest — but the listing copy should lead with those, not with "our website in an app".
The payments constraint, stated exactly (§0.2). Also relevant but not ours: Google's US alternative-billing programme starts charging a service fee on external payment routes from 1 Oct 2026 (10% on auto-renewing subscriptions) [reported] — US-only. In the UK the CMA designated Apple and Google with strategic market status on 22 Oct 2025 and opened a consultation on steering conduct requirements on 30 June 2026 [reported], so the UK rules on linking out are moving this year. Design for today's policy (no purchase inside the Android build); do not build on a loosening that has not happened.
What it costs us, honestly. One build session (Bubblewrap project, asset links file served from app.provenbatch.co.uk, store listing from the six G2 screenshots we already own), Dave's £20-ish and an ID check, plus review latency — and, if a personal account is used, a 12-tester closed test that the beta six plus friends could just about fill. Thereafter the listing is self-maintaining because the web app is the app.
Impact, argued. Two facts carry it. Our own competitor file records that review counts are tiny everywhere in this category (0–25) and "volume lives in app stores for consumer-style apps" — Bakesy has 1,074 Play ratings and 87 UK App Store ratings, Bake Diary was an App Store product [repo]. And no UK PPDS competitor is on Play at all: FoodCore, AllergenKit and Planglow are web products [repo]. That is an uncontested shelf in a place our phone-first buyer already searches. It will not be a flood; it is the only row in the table that keeps producing without anyone touching it.
2.2 Microsoft Store — free, tiny, do it while the package is warm
"With the new onboarding experience, there are no registration fees for either account type" [verified]; an individual account needs a government ID and a selfie, a company account a D-U-N-S number or business documents plus a work-domain email [verified]. PWABuilder submits the same package. The audience (Windows desktop) is close to irrelevant for a market trader, so this is worth exactly the 20 minutes it takes, for the listing and the link.
2.3 Apple App Store — defer, with the trigger written down
$99/year; individual enrolment shows the enrolee's personal legal name as the App Store seller; organisation enrolment requires a legal entity, a D-U-N-S number, binding authority, a work-domain email and a public website [verified]. Guideline 4.2 rejects apps that are "simply a web app wrapper with no native iOS features" [reported]. iOS users can already install the PWA from Safari [repo]. Trigger to revisit: Play produces measurable installs and either a legal entity (incorporation, on hold since 1 Oct 2026, #2980) has a D-U-N-S number in hand or the CMA's iOS interoperability commitments (consultation opened 10 Feb 2026 [reported]) improve web-app installability enough to make the wrapper unnecessary.
2.4 Integration marketplaces — all four fail, each for a different reason
Marketplace
What listing actually requires
Verdict
Xero App Store
From 2 Mar 2026: certification on every tier; a listing is optional only from Plus at $245 AUD/month, mandatory at Enterprise; Starter (free, 5 connections) and Core ($35 AUD, 50) cannot list; egress allowances with $2.40 AUD/GB overage [verified]. Plus onboarding 3 active connections in a 30-day review window [reported]
No. The listing costs more per month than three Pro subscriptions, before the integration exists
QuickBooks App Store
Free to list, but three sequential reviews — technical, then security, then marketing — with all critical/high/medium security findings remediated before publication, ~7 days for the initial security pass, and annual re-review [reported]
No for now. A permanent compliance obligation for a marketplace our users mostly are not in
Zapier
Public HTTPS API with public documentation, OAuth or API key auth, a non-expiring test account, English-only, "fully launched to the public"; free; 90 days in the directory with a Beta tag; partner tiers need 50 / 350 / 3,000 active users [verified/reported]
No. We have no public API [repo]; building one to earn a directory row is the tail wagging the dog
Shopify / Square
Shopify: build an app, pass review, 0% revenue share to $1M lifetime then 15% [reported]. Square: approval as an app partner before publishing [verified]
No. Both mean building a platform app for the minority of our buyers who sell through that platform; revisit only if #433's card-payment work shows a concentration
The pattern is worth keeping: integration marketplaces charge in engineering, and two of the four now also charge in rent. The one integration that would earn its keep is not a marketplace at all — Erudus, because it fixes onboarding as well as distribution [repo].
2.5 Software directories — the G2 family is the whole of the lane, and it is already packaged
Established 16 Sep 2026 and unchanged: Capterra is "powered by G2 Digital Markets" and its create a listing button goes to g2.com/products/new, so one G2 profile is the route into all four sites [repo]. #1944 holds it, deliberately gated until the public opening because G2 refuses products in alpha or beta, and it needs the marketing site redeployed that day so the served HTML no longer contains the "Private beta until 30 September" nodes [repo]. Nothing here is blocked on code; it is a submission only Dave can make, and the six screenshots, both logos, the banner and the field-by-field README are built.
Beyond that family: AlternativeTo takes free submissions (account must be 7 days old; an optional $5 buys a faster review) [reported] and is the one directory where our best-fitting query lives — the orphaned Bake Diary users, whose replacement pages every competitor wrote and we did not [repo]. SaaSHub is a form. Serchen/SourceForge stay optional exactly as §30 said: free and minutes, or not at all, because Google names low-quality directory links as spam. Product Hunt is free to post; realistic outcomes are 5,000–15,000 visitors for a top-3 finish, under 500 outside the top 10, and 50–300 signups counts as a great B2B launch [reported] — for a UK compliance tool aimed at market traders, the audience is wrong, so treat it as one afternoon for a permanent backlink, never as a launch plan.
2.6 Search plumbing — not content, and still not done
Search Console remains item 1 of §30's five, still owed: a Domain property verified by one DNS TXT on the apex, then the sitemap submitted once [repo]. Bing Webmaster Tools plus IndexNow is the same shape of job and was never on the list: free, an open protocol Bing/Yandex/Seznam share, a key file on the domain and a POST per changed URL [reported]. It matters more than Bing's share suggests because Bing feeds several AI assistants [reported], and our robots.txt is already fully open with the sitemap advertised [repo]. Neither task writes a word of content; both make the fourteen pages we already publish measurable and citable.
2.7 Member-benefit listings and affiliate links — the no-content half of the partnership plan
The 20 Sep pass ranked referrers by trust × reach × ease and put coaches first [repo] — but the ask it designed for them is a guest masterclass and a co-branded guide, which is content. Strip the content out and what remains is still substantial, because these bodies already publish partner lists with no software on them:
NMTF/deals — a bank, the AA, solicitors, an insurer, an accountant, no software; £150/yr membership; "actively seeks partnerships"; phone only [repo].
Encore Kitchens (25 sites, 550+ kitchens, "300+ brand partners") and Mission Kitchen (partner perks: Raja up to 50% off, Studio Mon 20%) — a perk line in a tenant pack [repo].
Farm Retail Association — a Farm Supplier Membership that buys a directory listing and a member discount, ~220 members in the finder [repo].
Accountants for food micro-businesses — Dead Simple Accounting publishes a tools list (FreeAgent, Xero, Mettle, Tide, SumUp); one email [repo].
Coaches, as an affiliate link rather than a session: the norms are CakeFlix 30% of the initial membership and Stocksmith 20% for 12 months [repo], and a code in Settings costs no tooling.
Each is one email or call, a promo code, and then silence. The risk is not cost, it is that a body with no published partner process may simply never reply — which is why this is ranked second on impact and flagged lumpy.
2.8 In-product distribution — the only compounding row
Two things, both already reasoned out and neither shipped:
The two-sided credit (referee's first paid month free, referrer gets a month when that payment clears, stacking to FreeAgent's cap), on the read-only banner so the archive state becomes a reactivation path [repo]. No referral SaaS: Rewardful/FirstPromoter/Tolt start at $49–69/month, a third of the whole budget [repo].
#688, the public QR allergen menu — open, Later, effort: high, security shape inherited from eho-view [repo]. It is a compliance feature first, but it is also the only surface where a ProvenBatch page reaches a stranger: a matrix on a stall, scanned by a customer, is distribution with a printed QR instead of a marketing budget. FoodCore gates its shareable matrix at £65; AllergenKit ships one at £9 [repo].
The same argument extends the EHO pack: shareable by link turns an inspection into an impression [repo].
3. Rejected, and why — so none of these gets re-proposed
§2.4 — rent, engineering, or an audience we do not have
Apple App Store, now
§2.3 — $99/yr, a personal-name seller line or a D-U-N-S, and a 4.2 fight for a surface Safari already installs
Trustpilot
§23's trigger (GA, ≥20 paying customers, a weekly review-answering commitment) is still unmet; free tier is not free in time
Google Business Profile
§23 — wrong surface for national software
AppSumo and lifetime-deal marketplaces
A one-off lifetime price at £9–39/month economics destroys the subscription it is meant to seed, and the buyers are worldwide software collectors, not UK PPDS businesses [inference]. Not researched further on purpose
Paid directory placement (Capterra/GetApp PPC)
~$2/click with ~$500/month floors against a £30–100 allowable CAC [repo]
§22/§30 — Google names them as link spam, and a product whose brand is the label proves what went in cannot look like a link scheme
Reseller / white-label / hardware lock-in
Planglow's software-only-with-our-labels model is the thing our positioning argues against [repo]; CompliChef's £310 printer bundle likewise
A fifth social channel
AGENTS.md #18 — final
4. What to do, in order
4a. October, and who owns each step
Dave, at the public opening (no date yet): redeploy the marketing site, then submit G2 (#1944) — the pack is built, the directory question gets answered from inside the account.
Dave, at the public opening (no date yet), 15 minutes: Search Console Domain property (one DNS TXT), then hand the token over; a session submits the sitemap and sets up Bing Webmaster + IndexNow.
A session, one sitting: the Play package — Bubblewrap project, Digital Asset Links served from app.provenbatch.co.uk, a store listing built from the G2 screenshots, and the no-purchase-path rule for the Android build. Dave opens the console account; the account type decides whether a 12-tester closed test is needed, so check that first.
A session, same sitting: submit the same package to the Microsoft Store (free).
Dave or a drafted email, October: NMTF, Encore, Mission Kitchen, FRA, Erudus, Brother, MUNBYN — one enquiry each, drafts-only through the outreach-upkeep rails, Dave sends.
A session, one sitting: the referral credit in Settings. #688 stays Later until its security addendum is reviewed.
20 minutes: AlternativeTo and SaaSHub rows; Product Hunt only if an afternoon is free.
4b. The one correction this pass makes to the 90-day plan
gtm-acquisition-research-2026-09.md §C11 is unchanged in substance, but three of its seven priorities are content channels (the coach masterclasses and podcast pitches at priority 2, the four new pages at priority 3, the Metricool queue at priority 4). If Dave wants the no-content half run on its own — a week of setup and then nothing to feed — it is: Play + Microsoft Store, the member-benefit emails, the referral credit, G2, the search plumbing, and the exact-match ads. That set has no weekly obligation attached to it at all, which is the whole point of the question.
4c. Issues that would need filing
None of this is filed except #1944 (G2) and #688 (QR menu). Candidates, not filed by this pass: Play + Microsoft Store distribution (one issue: package, asset links, the suppressed purchase path, both submissions), the two-sided referral credit, Bing Webmaster + IndexNow, and shareable EHO pack / matrix links.
5. What I could not verify
Whether the operator (David Biley trading as ProvenBatch, a sole trader; incorporation on hold since 1 Oct 2026) has a D-U-N-S number — it decides the Play organisation route. Apple organisation enrolment needs a legal entity (§2.3), so it waits for incorporation. Not in the repo; check with Dun & Bradstreet or Dave.
Whether a Play organisation account is genuinely exempt from the 12-tester closed test [reported only]. Establish inside the console.
Play Store search demand for our terms — Play's own search-volume data is not public, and no free tool was reachable that would answer it. The Bakesy rating counts are the only proxy.
Whether Capterra/GetApp/Software Advice inherit a G2 profile automatically — still #1944's open question, unchanged since 16 Sep.
FSB Marketplace listing terms for a supplier — FSB publishes member-benefit partnerships (the JournoLink offer, Aug 2026) but no public "list your product" route or price was found [verified absence within this pass]; the route, if any, is a phone call.
NMTF, FRA and National Craft Butchers partner prices — not published; a call each [repo].
Sources
Fetched and read 22 Sep 2026:support.google.com/googleplay/android-developer/answer/6112435 (US$25 one-time fee; personal vs organisation; identity verification; testing requirement for personal accounts created after 13 Nov 2023; Play Console mobile-app device verification) · support.google.com/googleplay/android-developer/answer/10281818 (Payments policy — cloud software and services; the "consumed outside… cannot be accessed in a Play-distributed app" exemption) · developer.chrome.com/docs/android/trusted-web-activity/ (Digital Asset Links, Add to Home Screen criteria, Bubblewrap) · developer.apple.com/support/enrollment/ (99 USD; individual seller name; organisation D-U-N-S, binding authority, work email, public website) · learn.microsoft.com/windows/apps/publish/partner-center/open-a-developer-account (no registration fees; individual ID + selfie; company D-U-N-S or documents) · developer.xero.com/pricing (the five tiers from 2 Mar 2026, App Store listing availability per tier, egress and overage) · docs.zapier.com/integrations/publish/integration-publishing-requirements (HTTPS, public launch, auth, test account, English).
Search extracts, not opened (all [reported]): Play 12-testers/14-days and the organisation exemption · policy 4.3 WebView rejection patterns and the TWA recommendation · Lighthouse ≥80 for TWA · Google's US alternative-billing service fees from 1 Oct 2026 · CMA SMS designations (22 Oct 2025), the Apple commitments consultation (10 Feb 2026) and the steering conduct-requirement consultation (30 Jun 2026) · Apple guideline 4.2 wrapper rejections · Xero's 3-active-connections review gate · Intuit's technical → security → marketing review sequence and annual re-review · Shopify's 0%-to-$1M-lifetime then 15% revenue share · Square app-partner approval · AlternativeTo submission rules and the $5 fast review · SaaSHub submission · Product Hunt traffic by rank and B2B signup outcomes · Bing Webmaster Tools / IndexNow mechanics and AI-assistant relevance · FSB membership pricing and the JournoLink member offer.
Pepesto — a supermarket price API, and where it fits ProvenBatch (researched 20 Sep 2026)
Researched 20 Sep 2026 from pepesto.com's own pages (home, pricing, recipe API, supermarkets, Tesco GB, Sainsbury's, API reference, MCP, agent-to-cart, shelf analytics, about, API terms, terms of use), the public openapi.yaml and the two other repositories under github.com/pepesto-solutions, plus company-register and search extracts for the corporate facts. Dave's brief: "this toolset looks like something that could be hugely valuable to make additions and enhancements to ProvenBatch — research this service in detail and come back with all the places you feel this could add real value to our users." Every ProvenBatch-side claim below was checked against the tree on main the same day (file and line given), not recalled.
How to read the grades.[verified] = the primary page, spec or repository file was fetched and read on 20 Sep 2026; [reported] = a search extract or secondary page not opened in full; [inference] = our reasoning from the evidence; [repo] = a fact from this repository or its recorded production measurements.
Decided the same day (Dave, 20 Sep 2026): record the research here and in RESEARCH.md §32, publish it in the admin research library, and file no issues yet. The value points in §6 are candidates, not a backlog.
0. Headlines
Pepesto is not a food-business tool, and its product data carries no ingredient list, no allergen declaration, no per-100g nutrition and no barcode [verified — the CatalogProduct and IndexedProduct schemas in openapi.yaml, and the Tesco GB and Sainsbury's pages, which list the fields returned]. It is a live supermarket price and catalogue API for 28 European chains, plus a consumer recipe-to-basket engine. The "nutrition data" on its homepage is four recipe-level integers from its own knowledge graph, not product data.
So it gives ProvenBatch's compliance core — declarations, allergens, labels — nothing, and the standing decisions on external label data (#432: declined permanently) and on nutrition (#28: no external lookup without a design record and Dave's word) are untouched by it.
Where it does add value is the money half: keeping a supermarket-bought pack cost true before the receipt arrives, pricing a shopping list at today's shelf price including offers, and removing the typing on the costing side of onboarding. Useful, not transformative. Ranked in §6, with the three that are worth an issue when Dave wants one.
UK coverage is Tesco, Sainsbury's, Asda, Morrisons, Waitrose and Ocado at roughly 4,000 cooking-ingredient SKUs a chain, refreshed daily with promotions [verified]. Aldi, Lidl, Costco and Booker are not covered [verified — absent from every chain list; Lidl is listed as "coming"]. Aldi is the first real receipt in RESEARCH.md §8a and "The Pantry" own-brand sits in live sourcings [repo], so any surface must name the chains it can price and stay silent elsewhere.
Three constraints shape any yes: a three-day limit on retaining their price history without a Partner agreement [verified, API terms]; a new sub-processor entry and a credential-inventory line before the first Deno.env.get() [repo, AGENTS.md #3]; and the vendor-risk posture — a 2023 two-founder Swiss startup with no disclosed funding, the same shape as Bake Diary (RESEARCH.md §11), so nothing stored in the app may depend on the API being alive.
1. What Pepesto is
Company. Pepesto Solutions GmbH, Adliswil, Switzerland, register CHE-299.368.059, founded 2023 [reported — Moneyhouse and Northdata extracts]. Two founders: Angel Dzhigarov (CTO; ex-Google, Knowledge Engine / Assistant / Bard) and Georgi Mitev (CPO; ex-DeepCode / Snyk / Sonar) [reported — LinkedIn and the about page]. No funding disclosed anywhere read. Their about page describes the mission as "redefining how families cook at home using generative AI and structured food intelligence" [verified] — the consumer product is a family meal-planning app (Flutter, Google Play), and the API is the B2B side.
Technology, in their words [verified, about page]: a fine-tuned LLM for meal planning and the shopping agent; a "proprietary food Knowledge Graph" spanning "80+ allergies, dietary restrictions, ingredient substitutions, and seasonal nutritional data"; 150,000+ licensed recipes mapped to knowledge-graph entities; 28 supermarkets in 13 countries indexed daily.
Products [verified, home and product pages]:
Product
What it is
Relevant to us?
Supermarkets API (/catalog, /products, /search, /promotions)
Daily product catalogues with live and promotional prices, one JSON schema across chains
~4,000 SKUs each for Tesco and Sainsbury's, "focused on the most common ingredients used for cooking" — "everything visible and related to food, cooking, and recipes in the [chain's] mobile web application". Refreshed daily; "promotional prices and multi-buy deals … are captured in each daily refresh". The other four UK chains were not read individually.
Ireland
Tesco Ireland, Dunnes Stores, SuperValu
Covers #827 on the day Ireland is switched on
Not covered
Aldi UK, Lidl UK ("coming soon"), Costco, Booker, Bidfood, Brakes, B&M, Amazon
Aldi is the founder's own first real receipt (§8a) and a live own-brand in sourcings [repo]
Eight currencies, eleven languages; 23 further chains "under evaluation".
2.2 What a product record carries [verified — openapi.yaml, components/schemas]
Absent from both: ingredients, allergens, nutrition per 100g, EAN/GTIN, brand as a field, country of origin, store or postcode availability. The Tesco GB and Sainsbury's pages confirm the same field list in prose [verified].
RetrievedRecipe (from /parse and /suggest): kg_token, title, ingredients[] (strings), steps[], image_url, nutrition{calories, carbohydrates_grams, protein_grams, fat_grams} (four integers), allergens[] (strings), portions. This is the only "nutrition" and the only "allergens" anywhere in the API, and it is a knowledge-graph estimate for a recipe.
The example in their own docs: Barilla Spaghettoni No. 7 1kg at Tesco — price: 340, currency: GBP, price_per_meausure_unit: "0.34 / 100g", quantity: {Unit: {HundredGrams: 10}, accurate_grams: 1000} — "same request, different country, identical schema" [verified].
2.3 Endpoints and per-call prices [verified — /pricing/, /ai-grocery-shopping-agent/, openapi.yaml]
Base https://s.pepesto.com/api/, Authorization: Bearer pep_sk_…. Prices in euros per request; Growth is half of Starter.
Endpoint
Purpose
Starter
Growth
POST /products
recipe_kg_tokens or manual_shopping_list + supermarket_domain → items[] of matched products with live prices, not_indexed_items, currency. Also takes as_of_date (YYYY-MM-DD) and user_id.
€0.04 with tokens · €0.32 with a manual list
half
POST /search + POST /retrieve
Index a product "beyond cached inventory"; async, search_session_id, optional webhook_url
€0.12 · retrieve free
€0.06
POST /catalog (preselected)
Re-price up to 50 product_urls in one call
€0.96
€0.48
POST /catalog (promo-only)
Every product on promotion at a chain
€3.20 (pricing page: Starter N/A)
€1.60
POST /catalog (full)
The whole chain index
€9.60 (pricing page: Starter N/A)
€4.80
POST /promotions
supermarket_domain → promoted products
as promo catalog
POST /parse
recipe_url or recipe_text (+ image via /oneshot), locale, generate_image → RetrievedRecipe
€0.32
€0.16
POST /suggest
query + personalization → up to 3 recipes from 1M+
Drives a checkout session turn by turn (load_page, run_js, screenshots)
free after a session
free
POST /link, POST /credits
Get the key after a Stripe purchase; check euro_cents balance
free
free
Plans [verified — /pricing/]: Starter = €29.90 credit packs, pay-as-you-go, "credits never expire", no minimum, no approval step (key issued after Stripe payment). Growth = €79/month with a 6-month commitment (€474 minimum), half-price requests, priority support. Partner = custom pricing, custom SLAs — and, per the API terms, the tier that unlocks price history and redistribution beyond the standard limits. "28 supermarkets and 13 countries included at every tier; no per-chain fees."
2.4 The public repositories [verified — github.com/pepesto-solutions]
openapi-spec — openapi.yaml, README with 14 public endpoints, common shapes (prices in minor units, kg_token handles, session_token references), example workflows. Updated 9 Sep 2026.
api-examples — "26 working projects", one per chain. The ones shaped like our value points: Plus NL price tracking with a webhook, Auchan PL weekly price tracking, Waitrose / Migros / Conad promotional scanners, Aldi CH "recipe cost optimization", Tesco IE "price gap analysis (Ireland vs GB)", Dunnes shopping-list re-ordering. Updated 17 Sep 2026.
pepesto-mcp — TypeScript, MIT. Tools: pepesto_oneshot, pepesto_predirect, pepesto_parse, pepesto_suggest, pepesto_products, pepesto_catalog, pepesto_credits. Run with PEPESTO_API_KEY=pep_sk_… npx -y @pepesto/pepesto-mcp. "The key is returned only once — store it immediately." Updated 23 Jun 2026.
3. The terms that shape any integration [verified — /api-terms/, /terms-of-use/]
Commercial use inside a paid product is "the core use case. You don't need separate permission to charge your users." You may "show prices, promotions, and product data to your end users" and combine it with your own content. Redistributing derived data — "shopping lists, meal plans, savings calculations, and similar" — inside a broader service is permitted.
No attribution required for live product and pricing data. (Recipes differ; irrelevant here.)
Caching:/parse up to 7 days; /products in step with their daily refresh; /catalog "cache aggressively, maximum one call per day per supermarket".
🔴 "Historical pricing: retain for up to 3 days only" without a Partner agreement. A price series built from their data is not ours to keep. One current observation per sourcing ("seen on tesco.com at £1.05 on 20 Sep") is defensible; a trend line is not, without asking in writing.
Accuracy and liability: "no product dataset is ever perfectly current"; they "cannot guarantee that the information is complete, up to date, or free of errors"; allergen filtering is "not a safety guarantee — always read the actual product packaging"; the service "does not provide medical, nutritional, or dietary advice"; liability for direct and indirect damages is excluded; downstream compliance is yours — "You are responsible for ensuring that your own use, display, redistribution, or commercial exploitation of the data complies with applicable laws", and if a retailer objects, "responsibility for that downstream use sits with you".
Suspension for breaching redistribution limits or threatening stability, "where circumstances allow, we'll get in touch first". EU law governs.
Rate limits are not stated.
Reading [inference]: the terms are written for consumer meal-planning apps and price-comparison tools, and they fit a costing feature comfortably. They say nothing about regulatory use because nothing in the data is regulatory. The retailer-objection clause matters: Pepesto's UK data is read from the retailers' own sites, and the risk of a retailer objecting to a price-comparison use is Pepesto's to manage first, then ours.
4. Vendor risk, stated once
A three-year-old GmbH with two founders, a consumer app, an API business launched into the "agentic commerce" wave, and no disclosed funding [reported]. Its pricing page changed shape at least once (the API reference and the pricing page disagree on which /catalog modes Starter can buy) [verified — both pages read the same day]. RESEARCH.md §11 records what this market remembers about a cheap tool disappearing (Bake Diary, May 2025). The consequence is a design rule, not a reason to say no: every surface built on Pepesto degrades to today's behaviour when the API is absent, and no stored figure in the app may depend on it. check-prices already has exactly that posture ("no price found on page") [repo].
5. What ProvenBatch already has on the money side
So that §6's value points are deltas against the shipped app, not wishes. Every line checked on main, 20 Sep 2026.
Costing is live from the primary pack, never stored.lineCost() → recipeComputed() → productUnitCost() (outputs/bakery-app-site/src/part-1.js:12279), from sourcings.pack_qty / pack_unit / pack_cost.
Price drift is reactive.priceDrift() (src/part-2.js:4468) compares a receipt line to the saved pack cost and offers Update pack cost — its own comment: "a silent stale pack cost is how a recipe quietly stops being profitable — this surfaces it at the one moment the true price is in front of you". Price history (supplierProductPriceHistory(), src/part-2.js:2926) and Insights' "If a price rise bites" are built from expenses. All of it fires after the money is spent.
check-prices is deployed and has no caller. The edge function (supabase/functions/check-prices/index.ts) fetches a supplier product page and extracts a price from schema.org JSON-LD, then price meta tags; SSRF-guarded (#886), quota'd, on both projects (production v19). No line of app code invokes it. The six UK supermarket sites are exactly the ones a page scraper cannot read reliably [inference].
Supplier Website URL lives on a sourcing (src/part-1.js:10355) and surfaces as the Buy link on a saved shopping list (USER-GUIDE.md "Where to buy"). When the user pastes a supermarket product page it is exactly Pepesto's product_urls key.
Shopping list, every plan. One row per ingredient, merged across products, rounded up to whole packs, grouped by supplier with subtotals, re-priced live from the saved pack cost (USER-GUIDE.md:1635, src/part-1.js:21626). Ends at "Save list, I'm off shopping"; "I'm back, check what I bought" follows.
Suppliers are already chain-aware.suppliers.chain links a tenant's supplier to the global chain_profiles registry (Tesco, Aldi, Asda, Sainsbury's, Costco, B&M, The Range, Amazon — outputs/migrations/496_chain_profiles_forward.sql). A supplier row knows it is Tesco: the natural map to Pepesto's supermarket_domain.
Multi-shop sourcings (#1983): one supplier product can be bought at several shops, each with its own pack cost; receiptSourcingFor() prices a receipt line against the right shop.
The real customer's shops [repo]: Tesco own-brand is the largest single brand among Whisk & Whimsy's confirmed sourcings (seven of twenty-one fit for the library, 28 Aug 2026 — decision 432); receipts observed from Aldi, Tesco, Asda, Sainsbury's and Costco (§8a). Pepesto covers three of those five.
Onboarding friction is measured (§17): a month after pack reading shipped, half the founder's own catalogue still carried the (enter declaration) placeholder; every hand-created supplier product also has to be typed a pack size and a price.
No nutrition at all (#28, Later, design-first; "no external lookup unless the design record and Dave say so"). parse-pack deliberately stops before the nutrition panel.
brand_products is "prefill — never a label input" (migration 553's own table comment; decision 432). Labels derive only from the tenant's own sourcing rows. Any external data sits on the prefill side of that line.
Recipe import already reads URL, photo, PDF and pasted text (parse-recipe, fetch-page) at Haiku cost, through the app's own parseRecipeLine.
Meal planning, menus and production forecasting are not built and were judged "not worth chasing" for the D-24 buyer (FoodCore gaps; see the master competitor sheet).
No competitor in §3 has live supermarket pricing. FoodCore's "supplier price history" is receipt-derived, as ours is. Craftybase/Stocksmith's "pricing guidance auto-updating when material prices change" (§24b) is the nearest shape, at $49–349/month for a different buyer.
6. Where Pepesto could add real value — ranked
A. Pack costs that stay true before you shop — highest value, cleanest fit
What. For any sourcing whose Supplier Website URL is on a covered chain — or whose supplier has chain set to one and whose name matches — fetch today's shelf price and show the same price up / price down → Update pack cost chip the receipt path already uses: on the supplier product drawer, the Ingredients list, and the Health check's "Quietly wrong figures" grade. Nothing is written without the tap.
Why it matters. It moves the drift check from "at the next receipt" to "before the batch is costed and before the sticker price is set" — the exact problem priceDrift() exists for, caught earlier. It is also the missing caller for check-prices: the contract ({items:[{id,url}]} → {results:[{id,price}]}) stays, with a Pepesto branch for covered hostnames in place of the scraper that never read tesco.com.
Cost shape./catalog preselected re-prices 50 URLs for €0.96 (€0.48 Growth). A weekly platform sweep of ≤500 tracked URLs ≈ €40/month on Starter; on demand ("Check today's prices" on a drawer or a saved list, metered through api_usage_log like every AI call) is pennies, and is the right pilot shape. A product outside the ~4,000-SKU index can be indexed once with /search.
Constraints to carry. Every figure says "seen on tesco.com, <date>"; one observation stored (observed_price, observed_at, observed_source), never a series (§3's three-day rule); the price reaches costing only through pack_cost via the user's tap, never a parallel cost path (HANDOVER: "summaryForRange is the only money figure; wrappers must not grow a parallel calc"); "no price found" degrades exactly as check-prices does today.
B. The shopping list priced at the shop, with offers and "cheaper elsewhere this week"
What. On a saved list grouped under a covered chain, price the rows at today's shelf price including the chain's promotions, show the trip total, flag rows on offer, and optionally price the same rows at the other covered chains ("£4.10 less at Asda this week").
Why it matters. The list already knows the packs and the quantities; this turns "estimated from my saved prices" into "what this trip costs today". Margins in this segment are thin and §24 shows pricing is the question these businesses actually ask. Nobody in §3 has it.
Cost shape. One /products call per list per chain (€0.32 with a manual list), or the A-path URL re-price where the packs carry URLs. Per tap, metered.
Constraints. Only the chain-grouped part of a list; an Aldi or Costco group shows nothing new and the screen says which chains it can price. Pepesto's consumer-grade name matching can pick a different pack from the user's usual, so a row must show the matched product name and pack size, never a bare number.
C. Onboarding prefill — pack size, price and product URL from the supermarket catalogue
What. In "+ new supplier product" and the pack-intake "New to you" path, when the supplier is a covered chain, a typed name offers Pepesto matches; picking one fills pack_qty, pack_unit, pack_cost, the Website URL and an image. Declaration and allergens still come from the pack photo — Pepesto has none, and decision 432 keeps labels on the tenant's own evidence.
Why it matters. It removes the typing on the costing half of every supermarket-bought record, which §17 shows is where onboarding stalls. The record gets a provenance stamp beside the existing "from a receipt / from a pack photo" ones.
Cost shape./search at €0.12 or /products per lookup; per tap, metered.
Constraints. Prefill only, never a declaration source; the same "observed <date>" marker the brand_products banner already uses (src/part-1.js:14702); no offer for a non-chain supplier; a pick list, never an auto-fill.
D. Promotions watch on Today — later, if A lands
/promotions per chain once a day at platform level, matched against every tenant's covered sourcings: "three things you buy are on offer at Tesco this week". A flat platform cost (six chains daily ≈ €290/month on Starter; weekly on Growth ≈ €40/month) that is only worth paying once A and B show people act on prices. Not for a pilot.
E. Ireland (#827) comes for free
Tesco Ireland, Dunnes and SuperValu are live, so A–C work in Ireland on the day it is switched on. Note only.
F. Shopping list → supermarket basket — an experiment, not a feature
/predirect is free and keyless: a ≤30-line list becomes a link that opens the Pepesto consumer app with the items matched at the chosen chain, and checkout happens inside Pepesto's app, not on the retailer's site with the business's own account, Clubcard or delivery slot. Cheap enough to try behind a flag; the consumer framing, the 30-line cap and the third-party hand-off make it a poor fit for a business buyer. Recorded, not recommended.
7. Where it does not fit — recorded so nobody rediscovers it
Nutrition (#28). Pepesto's nutrition is four recipe-level integers from its own knowledge graph: not per-100g values per ingredient, and not the seven-line UK declaration (energy kJ and kcal, fat, saturates, carbohydrate, sugars, protein, salt). Unusable for a prescribed-format panel, and #28 forbids an external lookup without a design record and Dave's word. If #28 ever chooses a lookup, the candidate is CoFID (McCance & Widdowson, free, UK government) — not this.
Declarations and allergens. No product-level ingredient or allergen data exists in the API. Even if it did, decision 432 declined external lookup for label content permanently, and their own terms call allergen filtering "not a safety guarantee". Nothing here may touch a label.
Recipe import.parse-recipe and fetch-page already read URL, photo, PDF and pasted text; /parse at €0.32 adds nothing a compliance user needs, and /suggest (1M recipes, €1.50) is a consumer meal-planning feature already judged not worth chasing.
Digital Shelf Analytics serves brands listed in supermarkets, not businesses that buy from them.
The MCP server and Agent to Cart are consumer and agent-builder products. The MCP could serve a session doing costing research for a guide; it is not a product feature.
8. The costs of saying yes — line items, not discoveries mid-PR
Sub-processor. A platform key means tenants never talk to Pepesto, but B and C send tenant content (ingredient names, product URLs) and A sends URLs the user pasted. That is a new row in privacy.md §7 and on /sub-processors (a Swiss entity; UK adequacy applies) — the same legal-adjacent work decision 432 priced for Open Food Facts. A credential-descriptions.yml entry for PEPESTO_API_KEY is required before the first Deno.env.get() (AGENTS.md #3), and it belongs as a Supabase function secret, never client-side.
Coverage honesty. Aldi, Lidl, Costco and Booker are not covered. Every surface names the chains it can price and stays silent elsewhere; nothing may read as "we checked and it has not changed".
Matching without a barcode. Pepesto carries no EAN, so the only exact join is the product URL; name matching is fuzzy and the matched product must always be shown. A is URL-keyed by design; C is a pick list.
Retention. One current observation per sourcing, never a Pepesto-derived series. Our own price history stays receipt-derived. Ask Pepesto in writing before any trend uses their data.
Spend. Pilot on Starter credits, per tap, metered through api_usage_log under the existing per-user caps and the shared spend ceiling; do not sign the six-month Growth commitment before a measured pilot. The fixed cost base is £50–70/month (§11); a €79 line would be the largest vendor after Supabase.
Degradation. Every surface falls back to today's behaviour with the API absent.
Sources
Fetched and read 20 Sep 2026 unless marked:
https://www.pepesto.com/ — product overview, chain list, currencies, languages
System-wide learning mechanisms — research & candidate portfolio
Research, 16 Aug 2026 · status at 16 Aug 2026 (later the same day): L1–L4 picked by Dave and planned in full — see system-wide-learning-plan.md and epic #551 (sub-issues #552–#556); L7 recorded there as parked, much later. L5/L6 not picked — re-propose from §6/§7 here if wanted. Originally: suggestions only. Follow-up to the receipt-parsing knowledge base (epic #493, all four phases live as of 15 Aug 2026), whose design record is receipt-parsing-knowledge-base-plan.md. Dave's brief, paraphrased: having built a cross-tenant learning mechanism for receipt parsing, what other system-wide learning mechanisms could be implemented to the advantage of our customers?
0. The answer in one screen
The codebase audit (16 Aug 2026, v0.198.3 tree) found exactly one learning loop in the whole system — the receipt one just built. Every other surface that generates correction or usage signal discards it. Ranked by customer value:
#
Mechanism
The wasted signal today
Customer win
Fit with ratified doctrine
L1
Supplier-product declaration knowledge base (parse-pack layer)
pack_photos.read_text vs confirmed_text — the correction delta is stored per tenant and never aggregated
Attacks the evidenced #1 adoption barrier (typing ~70 declarations by hand)
Strong — branded pack declarations are public-world facts, like chain receipt formats
L2
Allergen-term learning (the scanner's keyword lists)
Every auto-added-then-human-rejected allergen is logged per tenant (compliance_log) and never counted; false positives are found by hand and hardcoded
Deepens the stated moat (second-opinion allergen correctness)
Strongest — matched terms come from our own closed vocabulary, so zero tenant text crosses
L3
Recipe-import site knowledge (parse-recipe / URL scrape)
Extraction method + failure per domain goes to debug_logs and is never aggregated; AI-parse corrections discarded at confirm
Fewer "that page doesn't publish its recipe readably" dead ends
Strong — a recipe site's markup is a public-world fact; direct clone of chain_profiles
L4
Health-check defect prevalence
Per-check fire/accept/fix rates exist as data (dq_finding_acceptances.check_key was stored for exactly this) and are never aggregated
Better guidance copy + product priorities driven by what actually goes wrong
New-tenant ingredient creation is raw title-cased text; the fuzzy matcher has no synonym knowledge
Faster onboarding; better parse matching for everyone
Good, with a registry gate (unmatched names stay local)
L6
Food-safety template learning (#527 tie-in)
Which seeded checks tenants keep/rename/delete is never recorded
Starter checks and the SFBB outline converge on what real businesses use
Trivial — counters keyed to our own seed list
L7
Anonymised price benchmarks
Rich per-tenant price data (expenses, sourcing pack costs)
"What do others pay for flour?" — genuinely wanted
⚠️ Weak — amounts ARE the data; needs opt-in + volume. Recommend: park, evidence-gated
Recommended build order if pursued: L2 → L1 → L3, with L4 and L6 folded in cheaply along the way (they are one-migration counters). L5 is a curation project more than a code project. L7 is parked the way #493 Phase E is parked.
A general timing note that shapes all of this: at today's tenant count (~2 real businesses), the telemetry halves of these mechanisms have tiny N — but the curated-knowledge halves work at N=1. The receipt project's own lesson (and the July api_usage_log lesson: spend was invisible because the ledger didn't exist yet) is to lay the measurement rails early and cheaply, so the data is already flowing when the F&F trial multiplies the evidence stream, while the immediate customer value comes from curation.
1. What the audit established (the shared context)
Full survey run 16 Aug 2026. The facts that matter for every proposal below:
The ratified doctrine (13 Aug 2026, #493 §7/§12) generalises cleanly and should govern any new mechanism verbatim: derived knowledge only crosses tenants; a registry gates everything (unrecognised = stays tenant-local); human-curated, machine-measured (stats flow up automatically, knowledge flows down only through platform-admin edits); aggregates are the standing admin capability, row-level needs a #57 support session; telemetry is counters-only, default ON with a Settings opt-out honoured at write time.
The reusable machinery already exists: global reference tables with authenticated read and admin-RPC-only writes (allergens, mandatory_statements, chain_profiles); client-written counters-only event tables (receipt_parse_events); the opt-out pattern (business_settings.share_scan_stats, scanStatsOn() gate at index.html:14195, recordParseEvent at :14198); admin aggregate RPCs (admin_parse_stats); the audit-logged curation write (admin_upsert_chain_profile); the model-assisted drafting aid that saves nothing itself (admin-apidraft_layout). Each new mechanism is a re-instantiation, not an invention.
Three AI parse functions share one quota/spend ledger (parse-receipt, parse-recipe, parse-pack — 20 calls/user/day, £10/day ceiling). Only parse-receipt receives any knowledge; the other two run a fixed prompt for every caller.
Surfaces with no telemetry hook of any kind today: recipe parse/import, pack parse, allergen scanner outcomes, search, onboarding/demo uptake, pricing, food-safety check lifecycle.
2. L1 — Supplier-product declaration knowledge base (parse-pack)
Why first-rank. RESEARCH.md §2 records the evidenced biggest adoption barrier: entering ~70 ingredient declarations by hand — nobody types that to trial an app. Every mechanism that makes declaration capture faster and more accurate attacks the top of the funnel. And this is the surface where the storage half of a learning loop already exists and was designed for it: pack_photos keeps read_text and confirmed_text separately, with the migration comment (249) saying "a correction is the most interesting thing that can occur here". Nothing has ever read that delta.
What exists today.parse-pack (Opus, deliberately the strongest model) reads a pack photo → verbatim declaration + emphasised allergens + pack size, explicitly forbidden from inferring from the brand name. The client sends {image_base64, mime} only — no brand, no supplier, no hints. Brand is a free-text field with a tenant-scoped picker (brandOptions()). There is no global product identity anywhere; the dormant supplier_product_sourcings.barcode column (v0.75.0 removal) is deliberately unread.
The mechanism, in #493's three layers:
Measure (pack_parse_events, the #495 clone): one counters-only row per pack read — path (ai/ocr), legible, confidence, declaration edited y/n (the read/confirmed delta as a boolean/edit-distance bucket, never the text), allergens auto-added count, outcome. Gated by the same opt-out. Answers "how often does pack reading actually save typing?" — the adoption claim we currently cannot verify.
Knowledge base (brand_products, the #496 clone): a global, curated registry of branded supplier products — brand + product name + pack size + the verbatim declaration + verified allergen set + evidence status (observed/researched, same honesty chips as chain cards). A pack declaration printed on a national brand's packaging is a public-world fact on exactly the chain-receipt-format footing. Feeding: (a) prefill — when a tenant picks/types a brand+name that resolves to a registry entry, offer the known declaration and pack sizes ("from the ProvenBatch product library — check it against your pack", never silently applied, because the pack in their hand is authoritative and formulations change); (b) prompt hint — pass the candidate brand to parse-pack the way chain is passed to parse-receipt, with the same the-image-is-authoritative safety valve.
Curate (the #497 clone): admin stats (which brands are scanned most, edited most — counts only) + a registry editor. Seeding starts from Dave's/W&W's own confirmed packs with consent and from public product data, curated by hand; a draft_layout-style aid can draft an entry from a photo Dave supplies, human save + audit line always.
Cautions.
This is knowingly adjacent to the removed Open Food Facts/barcode tier (v0.75.0). The difference: OFF was a live third-party query in the scan path; this is our own small, curated, observed-evidence registry. But re-introducing any global product identity deserves an explicit decision from Dave with the v0.75.0 reasoning re-read first, and barcode re-entry (the natural join key — the column still exists) should be its own later decision, not smuggled in.
Formulation drift is a compliance hazard: a stale shared declaration confidently prefilled is worse than an empty box. The prefill must carry its observed date, never auto-apply, and the existing declaration-freshness machinery (#168 review months) applies to registry entries too.
Tenant declarations never enter the registry automatically — donation of a confirmed pack into the shared library is at most an explicit per-item opt-in (Phase-E-shaped, parked until the curated seed proves insufficient).
3. L2 — Allergen-term learning (the scanner's vocabulary)
Why it may deserve to go first despite L1's funnel value. The strategic read in RESEARCH.md §1: the moat is correctness and trust — "we check our own allergen detection". The scanner's knowledge is two hardcoded regex lists (ALLERGEN_TERMS, SAFE_COMPOUNDS, index.html:15570-91) whose every improvement to date came from a human noticing a false positive (the cocoa-butter → Milk defect is memorialised in a comment) and shipping a source edit. Meanwhile every tenant correction is already recorded — sourcing_allergens.verified flips, "unsupported" disagreement verdicts, compliance_logallergen_auto_added entries — and none of it is counted.
The privacy story is the cleanest of all seven, and worth spelling out because it enables sharing the term itself, not just counters: when the scanner fires, the matched term is a word from our own fixed vocabulary ("butter", "whey"), not tenant text. An event of the shape (term: butter, allergen: Milk, verdict: human-removed) contains zero tenant data by construction. The registry gate is the vocabulary itself.
The mechanism:
Vocabulary to the DB: allergen_terms global reference table (term, allergen, kind: match/safe-compound, evidence, notes) replacing the hardcoded regexes as master — exactly the SUPPLIER_RECEIPT_PROFILES → chain_profiles move, embedded copy kept as the offline floor with the same a-test-pins-the-floor discipline. Curation via an audit-logged admin RPC: a new false-positive suppressor ships by admin edit, not by app release.
Measure: allergen_scan_events — term, allergen, and what the human then did (kept / verified / removed / level-changed). Counters plus closed-vocabulary terms only; the declaration text never leaves (compliance_log keeps that, tenant-local, as today).
Curate: an admin view — "terms ranked by human-rejection rate" — which is literally the false-positive discovery process that today happens by accident, made systematic. A rejection spike on a term is the cue to research and curate a SAFE_COMPOUNDS entry every tenant benefits from at once.
Caution. This vocabulary is the compliance-critical path. The doctrine's human-in-the-loop rule matters most here: telemetry must never auto-edit the vocabulary, and (as with hints) an admin edit needs the audit line precisely because it changes every tenant's next scan. A missing allergen match is the dangerous direction — curation bias should favour adding match terms and be conservative about suppressors; the existing rule that auto-scan is additive-only and never removes a human-verified row stays untouched.
4. L3 — Recipe-import site knowledge (parse-recipe / URL scrape)
The most literal clone of #493, at low cost. Three import paths (URL scrape via fetch-page + three extraction methods; photo/paste via parse-recipe; on-device fallback) — and the correction surface (renderParsePreview) where users fix qty/unit/matches before saving. All corrections discarded; the only per-site signal is a debug_logs line that already writes the raw URL platform-wide (arguably a worse privacy posture than an aggregated, registry-gated mechanism would be — a URL can identify a person's blog).
Measure (recipe_parse_events): path (url/photo/paste/csv), extraction method (json-ld/wprm/microdata/ai/manual), lines parsed/edited/matched/added-as-new, confidence, outcome. Registry-gated site field: canonical name for recognised public recipe sites, NULL otherwise — never the raw domain (the chain-gate rule transplanted). The importUrl dlog should stop recording raw URLs in the same pass.
Knowledge base (site_profiles): per-site facts — which extraction method works, known markup quirks, unit conventions (US cups vs metric — feeds the fraction/unit normalisation the prompt already does), a prompt-hint block for the AI path. A recipe site's markup is a public-world fact; seeding starts from the sites the extractor already special-cases (WPRM) plus the UK-popular set (BBC Good Food etc.), evidence-labelled.
Curate: same admin card pattern; "sites ranked by failure rate" plus the bare (unrecognised) bucket count showing how much import happens off-registry.
Customer win: fewer failed imports, better first-parse accuracy, and an evidence base for which sites to support next — today that priority call has literally no data behind it.
dq_finding_acceptances stores check_key per accepted finding, and the migration comment says why: "for an admin view across fingerprints" — the aggregate view was anticipated and never built. dq_weekly_snapshots deliberately stores only the three class totals, not which of the 15 checks fired.
Measure: widen the weekly snapshot's counts jsonb to per-check-key counts (our own closed key set; zero tenant content). One client-side change + no new table.
Curate/act: an admin aggregate RPC — fire rate, acceptance rate, and time-to-clear per check across tenants. This is product learning more than parser learning: it answers "which defect classes do real businesses actually hit and which do they accept as noise" — directly steering guidance copy, Health-check explanations (the exact thing W&W's feedback asked for), and backlog priority. A check everyone accepts-as-noise is miscalibrated; a check that fires everywhere is an onboarding gap.
No new privacy surface: counts of our own check keys, same opt-out, same aggregate doctrine.
RESEARCH.md §9 already cites the USDA FoodData Central model: a Foundation (generic) tier and a Branded tier. L1 is the branded tier; this is the generic one. Today a new tenant starts with zero ingredients, no autocomplete, no synonyms; bestIngredientMatch is a bigram similarity with no knowledge that "caster sugar" and "castor sugar" are one thing; receipt-line matching strips store brands via a hardcoded regex.
Knowledge base (ingredient_vocabulary): curated generic ingredient names + synonyms + category, seeded from public reference lists — powering "+ new ingredient" autocomplete, synonym-aware matching in all three parsers, and optionally a "starter ingredient set for your business type" at onboarding (the business_type field exists and is unused for tailoring). Works at N=1 tenants; pure curation.
Measure (optional, later): counts of created-ingredient names that match the vocabulary (registry gate — an unmatched name may be "Mill Lane Flour Co." and stays local; only a bare unmatched count crosses). Tells curation what the vocabulary is missing without reading tenant text.
Honest caveat: the seed list is the work here, and it needs the D-24 discipline (any small food business — a butcher's and a jam maker's vocabulary, not a baker's).
SAFETY_SEED_CHECKS seeds six starter checks once, names-only by decision; the SFBB-outline seed (research recommendation D1) is planned in epic #527. Nothing records which seeds tenants keep, rename, delete, or add. A counters-only event at check create/rename/delete/complete — keyed to the seed list (closed vocabulary; renamed/custom checks cross as a bare "custom" count) — would let the starter set and the #527 SFBB outline converge on what real businesses actually run, which is precisely the evidence W&W's feedback ("prebuilt amendable clean checklists") says they want the seeds to be. Build it with #527's wave rather than as its own project.
8. L7 — Anonymised price benchmarks: attractive, and recommended PARKED
The one customers would most readily name if asked ("what do others pay for flour?"), and the one that most stresses the doctrine: #493's boundary was "amounts: never" — here the amount is the product. It would need: explicit opt-in (not opt-out — a different legitimate-interest posture), double registry gating (chain-linked supplier products and vocabulary-matched generics only), suppression thresholds (at <N businesses per cell nothing shows), and enough tenants that a benchmark isn't "Dave + W&W with a mask on" — the small-count caution in #493 §8 is fatal here today, not just cautionary. Also, half the value already ships without any sharing: the tenant-local price-trend work (#439 sparklines, priceDrift, Insights price changes) is live.
Recommendation: park it the way Phase E is parked — written down, no issue, gated on real tenant volume (say ≥20 active businesses) and on L5's vocabulary existing (benchmarks need the generic tier to compare like with like).
9. Adjacent but deliberately out of this portfolio
Per-tenant learned corrections ("this business always corrects X→Y") — #493 §10 already flagged it as worth its own backlog item: tenant-local, no privacy question, no admin involvement. Not system-wide learning; listed here only so it isn't forgotten.
Search zero-result telemetry — a counters-only zero-result rate per segment is nearly free but the lookup screen searches only the tenant's own data, so the learnable is thin. Fold into L4's snapshot widening if ever wanted; not worth its own mechanism.
Auto-learning of any kind (telemetry mutating prompts/vocabularies without a human) — re-rejected for every mechanism above, same reasoning as #493 §10: unauditable drift in shared compliance-adjacent behaviour.
10. Cross-cutting build notes (apply to whichever are picked)
One consent story, not seven.share_scan_stats and its Settings copy are receipt-specific. Before a second telemetry stream lands, decide: widen to one "share anonymous usage statistics" umbrella (one toggle, privacy-notice clause rewritten once — and the #211 legal stack is conveniently still pre-solicitor-review), or per-feature toggles. Recommendation: one umbrella toggle + one clause listing the streams, each stream honouring it at write time; per-feature toggles multiply Settings noise for no real control gain.
Generalise the client pattern once: recordParseEvent's shape (opt-out gate, fire-and-forget insert, counters-only row) becomes a shared recordEvent(table, row) helper rather than three copies.
Every new enum/CHECK pair registers with drift_info() in the same migration (standing rule); every new tenant event table joins the #209 export census and #210 purge script; global reference tables join the reference_tables export list.
Admin console cards follow the a1.5.0 pattern (aggregate RPC + evidence chips + audit- logged curation writes); the draft_layout drafting-aid shape reuses for pack/site/vocabulary curation aids — draft-for-review from Dave's own material only, human save always.
Suppression thresholds before open sign-up (the #493 §8 caution) apply to every aggregate RPC above, not just receipts — worth building the threshold into the first new RPC and retrofitting admin_parse_stats at the same time.
11. Suggested next step
Nothing here is decided. The #493 precedent worked well: Dave picks the mechanisms worth pursuing → a full plan doc with a §7-style decision block per mechanism → ratification → issues. Suggested picks if the portfolio is thinned to three: L2 (moat, cleanest privacy, smallest), L1 (the adoption barrier, biggest customer win, needs the OFF-history decision), L4 (near-free, feeds product steering immediately). L3 rides whenever recipe import next gets attention; L6 rides with #527; L5 is a curation backlog; L7 stays parked.
Notes
Last updated 6 Aug 2026
UI polish research (#388)
Committed 6 Aug 2026 — the permanent copy of the research comment posted to #388. The issue body carried the condensed findings; this is the complete record, so it's findable from the repo rather than only from a GitHub comment.
The consistent recommendation from NN/g and Baymard-derived guidance: help text belongs below the field, visible, when the field is genuinely complex — users need continuous reference while typing.
The equally consistent counterweight (GOV.UK): hint text should exist only after you've tried improving the question, and should be one short sentence. Content hidden behind a disclosure ("details" components, tooltips) is missed by a large share of users — never hide something the majority needs.
Polaris's rule: help text is one short line, used sparingly, "link out if more is needed."
Implication for us: our compliance hints fail the length rule, not the visibility rule. The fix is a copy rewrite to one line, with elaboration behind the ⓘ — not full relocation.
Convention across Polaris/Material and surveyed admin products: the info affordance rides with the field label, small, immediately after the label text. #356 placed it below the input — nonstandard, and part of why it reads oddly.
Hover doesn't exist on touch. Tap-triggered, and it must dismiss on any outside tap.
Passive info conventionally wants a popover (non-modal, anchored) rather than a modal — but at 390px, anchored popovers and small centred cards converge, and the anti-clipping argument against floating anchored popovers on this app's surfaces stands. A small centred card on a scrim, tap-anywhere/Esc dismiss, no buttons is the right mobile-first shape, and reusing the existing uiModal machinery inherits its focus-trap and focus-restore for free.
Tooltips must not contain interactive content (a11y rule) — ours won't. aria-describedby association is kept by leaving the full hint text in the DOM visually hidden.
Colour: severity colour conventions (Carbon status indicators, Material alert severities) are universal: amber/red = something is wrong here now. A coloured ⓘ on a healthy field would read as an error state and dilute the app's real warnings. Verdict: neutral ⓘ everywhere; importance is carried by the visible one-liner, not icon colour.
The underlying rule every system agrees on: mark the minority. GOV.UK/Polaris say "(optional)" because in their domains most fields are required. NN/g leans asterisk-on-required. In ProvenBatch most fields are optional, so the minority is the required set — the "(optional)" parentheticals are the rule applied backwards.
Decision (Dave, 6 Aug): delete every "(optional)"; mark the few required fields with one consistent indicator wired to aria-required. Save-time validation naming the field stays as the backstop.
Baymard's testing shows width-mismatched fields cause measurable hesitation — users pause, re-read the label, sometimes type extra characters and delete them. Fixed-length inputs (a 4-digit quantity, a percentage) should be sized to their content; variable-length text gets the "normal" width.
Mobile form consensus: strict single column, labels above fields, ≥44px touch targets, one consistent spacing scale. Multi-column rows that wrap unpredictably at narrow widths are the main source of the "misaligned boxes" effect.
Implication: width tokens by content type (xs: qty/%, sm: money/dates/units, full: names/text) applied across all field sites, plus a drawer-wide single-column mobile rule.
Scope addition (Dave, 6 Aug 2026): currency display at two decimal places
All currency fields must display with two decimal places consistently — before this batch they didn't always (e.g. a price entered as "5" stayed "5"). Added to the batch scope:
Displays: audit every money figure to confirm it routes through the shared gbp() formatter (which is where 2dp lives) — any raw-number money display is a bug.
Inputs: money inputs format to 2dp on blur (entering "5" becomes "5.00"), without fighting the user mid-typing, and without changing what's stored (storage stays full-precision; this is display only).
Test coverage: polish388.js asserts no money input lacks the blur formatter and no money display bypasses gbp().
What shipped
Both the placement fix (ⓘ after the label, centred dismissible card, no colour-coded severity) and the required/optional rework (mark the minority) landed as designed. See outputs/DESIGN-LANGUAGE.md's "Helper text is tiered", "Field width matches content, not the grid", "Currency inputs are always 2dp", and "Drawer forms are single-column on phones" sections for the pattern as built, and outputs/tests/hint356.js + outputs/tests/polish388.js for what's enforced going forward.