#714 — the whole-business glance: one tile per area, derived from outputs/ops/jobs.json's registry and this project's own ops_runs rows. Amber is a job whose latest run failed; a tile opens that area's Activity. Decisions are asked in the Stella bots' chats now, not here (#3181).
The registry: every repeatable business task, what runs it, how often, and what it hands back to you. Source: outputs/ops/jobs.json — outputs/docs/ops/jobs.md is the prose companion.
| Job | Area | Runner | Cadence | What you get back | Last outcome |
|---|
| Fired | Job | Outcome | Summary |
|---|
#733 — every request from provenbatch.co.uk/beta, plus anyone added by hand. Inviting mints that person their own sign-up code (#1034 — one code, one person, max_uses = 1, nothing typed) and emails them straight away to say they have a place and when the beta starts; when that business signs up carrying the code, the row flips to Signed up and points at the real business on its own. #741 — the invitation itself auto-sends on each row's send date (1 Sep 2026 by default), and everyone who fills the form in is thanked within the minute. The ✉ marks below say what has actually gone out. #910 — the page also has a one-field "just want updates?" box; those people are counted separately as updates-only, kept out of the stage counts and the decision list, and are never invited to anything — the mailer filters them out of all three emails, so it holds even if a status is set by hand. Move to beta is there for anyone who later asks for a place. #1594 — Marketing list is the standing roster of everyone who would get the next campaign, with who opted in, who unsubscribed, and which updates they have actually received — not only the names Preview shows at send time.
A partner code is minted here, with no trading proof. It is single-use, takes one of the 30 places and one of the partner places, and whoever signs up with it is a partner. The partner places only stop new partner codes: one already handed out always works.
| Code | For | Status |
|---|
| Asked | Who | What they make | How they label | Heard via | Status | Emails |
|---|
For an enquiry that arrived by email, at a show, or through an introduction — so the pipeline is the whole picture, not just the people who found the form.
#1059 — a progress update to everyone who ticked "keep me posted", written here rather than in the repo and sent from here rather than from GitHub. Preview renders the email exactly as it will arrive and names every person who would receive it; it sends nothing and records nothing. The count also opens the Marketing list (#1594), which is the standing roster rather than a send-time snapshot. A sent campaign's recipient count drills into who actually holds a ledger row for it. Send is only offered from that preview, restates the subject and the count, and refuses if the list has moved since you looked. Schedule… sits beside it (#2443): pick a UK day and a quarter of an hour, and the same send runs then, with the same count check. Editing the copy before then unschedules it, and Unschedule… on the row stops it. One email per person per campaign — press Send twice and the second press sends nothing — and every message carries a working unsubscribe link. Write plain paragraphs separated by blank lines; [text](https://…) becomes a link and everything else is shown as typed.
| Campaign | Subject | Status | Sent | Recipients |
|---|
Start a time-boxed, reasoned session to view one business read-only (#57/D-14) — capped at 60 minutes, and visible to that business's owner in their own Settings. Then use "Open app" above to act as them.
| Name | Category | Plan | State | Subscription | Trial | Members | Last active | Created | Source | Attribution | Registration |
|---|
The create-account card's optional "How did you hear about us?" answer. "— not said —" is not a channel, just the count still to close.
| Source | Businesses |
|---|
Percentage discounts on our subscription, minted one per customer from the Discount… button on the row above. Each one is a Stripe coupon plus a promotion code bound to that customer, single-use and dated — so it cannot be forwarded, cannot be used twice, and stops working on its own. The customer types it into the box on Stripe's checkout page and sees it in their own Settings. Redeemed is written by the Stripe webhook when the discounted subscription starts, never from here; Expired is worked out from the date rather than stored, so nothing has to sweep it. Revoke kills a code that has not been used yet — a discount already attached to a live subscription is a billing change, made in Stripe, with the customer told.
| Customer | Code | Off | For how long | Expires | State | Why |
|---|
Which nations the product is switched on for. Enabling a row starts that market immediately: its label rows, food-safety pack, tax profile and currency. Ireland ships off — turning it on is the market-start act, not a deploy. This console talks to the environment in the address bar, so staging flips staging and production flips production.
| Nation | State | Currency | Pack | Regulator |
|---|
Who sends us businesses, and what they have sent. An affiliate earns a share of what their referrals pay, worked out here and paid by hand. A member perk gives a club's members an offer through a sign-up code. A listing is a directory or marketplace page. Every count is read from the businesses themselves: the aff code their trial link carried, the sign-up code they joined with, or a credit added by hand below. Paying is an active subscription on a customer account, the same rule as Revenue. Est. commission owed is one month at list price for each paying referral. The date a business first paid is not recorded yet, so stopping when an affiliate's months run out is still yours to do. Partners never emails anyone, pays anyone or changes a price.
| Partner | Type | Status | Offer | Share link | Referred | Last contact | Notes |
|---|
For a business a partner sent that signed up without their link or code. A credit counts exactly as a code would. A business that did use a partner's code is credited by the code, so it cannot count twice. Removing a credit takes it out of these counts and changes nothing else.
| Business | Partner | Credited | Note |
|---|
Give a month, get a month: one business's ?ref= link bringing another. The credit job issues both sides 14 days after the new business's first payment, with no step here. This card is for the exceptions. Approve clears a flagged or fraud-rejected referral, and the job then issues it on its next run (the 12-a-year cap still applies). Reject stops one the job has not issued yet. Reverse takes back the unused part of one Stripe credit; anything the customer has already spent is written off, never clawed back. Every action needs a reason and is recorded with it.
| Referral | State | New business's month | Referrer's month |
|---|
Active customers only — the same rule as the Revenue card (#1592). MRR is estimated from plan prices (Starter £9, Standard £19, Pro £39) and assumes monthly billing. ARR is 12 × that MRR, not a separately invoiced annual figure. Friends, Beta and Internal never count, even when Stripe says active.
platform_stats_dailyNightly estimated MRR, Active-only since #1592. Same snapshots the Platform Trends card reads; this view is the money line rather than the 14-day operating table.
| Date | Est. MRR | Active in 7d | Businesses |
|---|
Active paying businesses by plan, and ARPU as Active MRR ÷ that headcount. A business on trial with no plan does not appear here.
| Plan | Businesses | Est. MRR |
|---|
SaaS cards we will not invent numbers for. Each one is a child spike once the input exists — not a gap on this page.
| Signal | Why it is deferred |
|---|---|
| CAC | No ad-spend ledger and no attribution pipeline. Signup source (#1871) is a self-reported channel, not cost per acquired customer. |
| LTV | No cohort survival. Nightly MRR snapshots are a stock, not a customer lifetime. |
| GRR / NRR | Need expansion, contraction and churn in pounds for the same cohort across a period. We snapshot total MRR, not per-customer movement. |
#878 — how long each Tier-1 journey actually takes a customer, and whether that is moving. Field figures are computed from journey_events marks, consented businesses only. Active time sums foreground gaps capped at 120 s, so a batch interrupted by a trip to the cash and carry stays comparable to one that was not — p50 and wall time are in the drawer, because p75 is the figure we alert on and a row carrying four different times gets read as none. Δ compares the last 7 days' p75 against the prior 28. Lab is the deterministic harness run: taps first, because a tap count is noise-free and a millisecond count is not. Click a row for the detail drawer.
| Journey | Started | Completed | p75 active | Δ p75 | p75, 14 days | Worst drop-off | Lab |
|---|
J01 Sign up is a milestone funnel, not a timed session — it spans days, so it reports its completion against the business's first ever printed label and never an active-time figure. Timing it as one sitting would invent a number.
The same outcome by different routes — the card that answers "is the shelf sweep actually faster than adding one by hand?", asserted in the user guide today and measured nowhere. Bars are median active time, scaled to the slowest route. Taps is the structural cost from the journey file, which moves only when the design moves. A route used by fewer than learning_suppression_n() distinct businesses is suppressed: a median over three tenants is a number about three tenants.
The deterministic harness run beside what customers actually experience. They measure different things and are never averaged together — the point of showing them side by side is that the lab moves first. A lab-only rise is a regression catchable before anyone feels it; a field-only rise is something the harness cannot see, which is usually data volume or a slow connection rather than the code. Both are indexed to the earliest release shown = 100 so they can honestly share one axis — a few seconds of machine time and a several-minute human p75 are different magnitudes, and two bars each scaled to their own maximum is a dual axis in disguise. The raw values are in the table beneath.
#470 — how far a business gets, and how long it takes. Derived from the tenant tables themselves, never from instrumented creation paths, so a new creation path cannot silently drop out of it. Reached is "ever got here"; Timed is "and we know when" — rows created before 25 Aug 2026 have no timestamp, so Timed is the smaller number and the median is computed over it alone. Businesses that turned off usage statistics are excluded.
| Step | Milestone | Reached | Timed | Median days from signup |
|---|
Is anything we shipped actually being used. Event rows come from activation_events (surfaces that leave no other trace); derived rows are counted from the tables that already record the act. Last touched is the number that matters — a healthy lifetime count with an old last touch is a dead feature. roadmap_expanded is #708 stage 2's gate metric.
| Surface | Source | Businesses | Touches | Last touched |
|---|
#500 — the founder-operating question: who's actually using the product, how much. #1589 — supplier-product and ingredient counts sit here with the other catalogue-size figures, not on Businesses. #2887 — Feedback counts the in-app feedback a business has sent, all time (feedback_items); feedback sent by email isn't counted. AI spend is per-user (api_usage_log has no business_id), joined through membership — a member of two businesses attributes to both.
| Business | Products | Supplier products | Ingredients | Recipes | Orders | Batches | Labels | Feedback | Storage | AI calls | AI spend (total / 30d) |
|---|
#721 — the abuse circuit breakers hard-coded in parse-pack/parse-receipt/parse-recipe/provenbot, made visible and editable here. A change reaches the functions on their next call — no redeploy. Shared ceiling is the combined daily Anthropic spend across those four. #752 — per-business ceiling is a second, tighter gate on the same spend, so one busy tenant can no longer exhaust the shared ceiling for everyone; both apply, and unattributable spend counts against the shared one only. #1939 retired transcribe — four rows, not five.
| Function | Spend today | Calls today | Daily cap | Shipped default |
|---|
#1047 — what Anthropic has actually billed, from the Usage & Cost Admin API, read once a night by anthropic-cost.yml. The card above it is our own estimate, computed per request from the prices the edge functions carry; that is the right figure for a circuit breaker and the wrong one for a bill. Anthropic exposes no credit-balance endpoint at any price — not in the Admin API, and the Spend Limits API that would come closest is Claude Enterprise only. Remaining credit below is therefore derived: the top-ups you have recorded here, minus authoritative spend since the first of them. It is only as right as those entries. Spend is grouped by workspace so a staging figure that runs away is distinguishable from a production one; Anthropic reports the Default workspace as null and it is shown as Default. Priority Tier spend is excluded by the cost endpoint itself, so it is missing here too, and today's figure is always partial — the day is not over.
| Workspace | Today (partial) | 7 days | 30 days | Tokens (30d) | Last seen |
|---|
The quietly valuable half. If api_usage_log's prices go stale — a model repriced, a new model in service — this is where it shows up, as a widening gap. Compared in USD: our estimate is stored in GBP having already been multiplied by the USD_TO_GBP constant those functions carry, so dividing it back out compares like with like and leaves the gap measuring token prices rather than the exchange rate. The two sides are not scoped alike, and cannot yet be. Anthropic's figure is organisation-wide — every workspace. Ours is this environment's api_usage_log only, covering every slug that calls Anthropic (the four edge functions plus admin-api's two). Scoping Anthropic's side to match would need a per-workspace split that does not exist yet: both Supabase projects currently share one inference key, so all spend lands in one workspace. #1051 separates them, and this line becomes like-for-like once it does. Until then a standing difference in that direction is expected and is not a stale price. One known gap: the scheduled Actions scripts that call Anthropic (social drafting, recall watch) write no api_usage_log row, so their spend appears on Anthropic's side and not on ours. A small standing gap in that direction is expected and is not a stale price.
| Day | Anthropic actual | Our estimate | Gap |
|---|
Anthropic will not tell us what you paid in, so this is the only way the balance above can exist. Enter each top-up as it happens, in US dollars — the currency Anthropic bills in.
Derived from business_settings (#500/#492 §6.1) — Stripe fields plus the #474/#662 cohort, no live Stripe call. Status is Active / Beta / Friends (Internal is listed separately and never counted as Active). Estimated MRR assumes monthly billing and is Active only: an annual subscriber inflates its row by up to ~12×.
| Plan | Status | Businesses | Est. MRR |
|---|
| Date | Businesses | Active in 7d | AI spend (7d) | Est. MRR | Disk |
|---|
Each place Jev makes a decision for a customer runs as its own pilot. It starts in shadow: Jev's answer is logged next to today's rule and the choice the person actually makes, and nothing on their screen changes. It is switched on only after its readout clears the criteria agreed before the run. Jev at ≥ 0.9 is how often Jev is right when it is at least 90% sure, the confidence at which it would fill a field in for the person. Today's rule is the current rule scored on the same rows. Only rows carrying the person's final choice are scored. Open a row to see that pilot's measures.
| Pilot | Stage | Evidence | Jev at ≥ 0.9 | Today's rule | Guardrails | Verdict | Open |
|---|---|---|---|---|---|---|---|
| Loading… | |||||||
Every pilot calls the one system-one function (#2236). Its budget sits outside the Anthropic ceiling on purpose, so a Jev fault can never use up the budget for reading receipts and packs. Calls and spend are read from api_usage_log, which records successful calls only. Fallbacks, waits and model times come from system_one_calls (#2408), which records every call's outcome and how long it took. A fallback is a timeout or an error, counted as a share of the calls that reached Jev; a call turned away at the cap is not one. The model time is the model call alone, the part the 2 s timeout bounds, and that is what its guardrail judges (#3152). A day is judged only on at least 10 timed calls, so one cold start on a quiet day cannot hold a pilot. The caller's whole wait, from the sign-in check until the answer is recorded, is shown beside it and not judged. It is timed inside system-one, so it leaves out the network and any start-up time, and the app's real wait is a little longer. Calls recorded before #3152 carry no model time and read not measured. On a database without #2408's migration, fallbacks and waits read Not recorded; without #3152's, model time does.
#497 (#493 Phase C) — cross-tenant scan-quality aggregates from receipt_parse_events (#495's counters-only telemetry). Aggregates are the standing capability; anything row-level about one business's scans needs a #57 support session (ratified 13 Aug 2026). The (unrecognised) bucket is a bare count — no merchant names exist to show.
| Chain | Scans | AI / fallback | Confirmed | Failed | Confidence | Edited | Unmatched | Ref read | Totals agree | Businesses |
|---|
Spend is api_usage_log's parse-receipt ledger — on a screen for the first time. Quality and money deliberately live in different tables (ratified decision #2).
| Day | Scans | AI | Fallback | Failed | AI spend |
|---|
#497/#499 — the chain_profiles registry (#496): the master for parse-receipt's prompt hints and the offline fallback's fingerprints. Editing a chain's hints changes what every tenant's next scan is told — every save is audit-logged. Format cards render live from each chain's layout descriptor, so they cannot drift from what the parser is actually told (#499); every example in a card is fabricated, never a real receipt's contents.
| Chain | Kind | Evidence | Aliases | Updated |
|---|
#553 (#551 M1) — the brand_products registry: a small, hand-curated, evidence-dated library of national-brand declarations. Offer-only — the sourcing drawer's prefill banner never auto-applies it, and labels derive solely from a tenant's own sourcing rows; this library only speeds data entry. pack_parse_events aggregates below never show a brand string, by design — brand_matched is a bare boolean.
| Path | Scans | Confirmed | Discarded | Failed | Brand matched | Name read | Size read | Edited | Confidence | Businesses |
|---|
Answers #768: how often the picker's top fuzzy name match is actually the one the human picks, so a "give a strong name match the barcode's named button" decision has real evidence behind it instead of a guessed threshold. A barcode match is always used_top by construction — read its row as a baseline, not something to improve on. used_top right after a confirmed outcome is slightly optimistic: a human can tap the top match and only discover it was wrong at the declaration-check step, which records no correction here. No name, brand or declaration text is ever in this table — only the verdict shown, what the human then did, and the top match's evidence/similarity/margin.
| Verdict shown | Reads | Used top | Used other | Used list | Created new | Mean evidence | Mean similarity | Mean margin | Businesses |
|---|
#675 — every tenant-accepted pack-intake read that carried a brand and name lands here, never straight into the library. A draft is only a prefill; nothing is saved until you review it in the editor and press Save there.
| Brand | Product | Barcode | Declaration | Received |
|---|
| Brand | Product | Pack sizes | Declaration | Evidence | Updated |
|---|
#555 (#551 M3) — the site_profiles registry: a small, hand-curated knowledge base of recipe sites (preferred extraction method, unit convention, format hints). A domain not in this registry never appears anywhere in recipe_parse_events, even hashed — the (unrecognised) row below is the curation queue, showing how much import happens off-registry.
| Site | Imports | Confirmed | Discarded | Failed | Edited | Mean lines | Confidence | Businesses |
|---|
| Site | Host patterns | Preferred method | Unit convention | Evidence | Updated |
|---|
#554 (#551 M2) — the allergen_terms registry: the scanner's vocabulary, curated here instead of by app release. Editing a term changes what every tenant's next allergen scan finds — every save is audit-logged. Asymmetric conservatism: a new match term may be added on modest evidence (a false positive is annoying); a new suppressor needs strong evidence (a false negative is dangerous — never add one on a hunch).
Ranked by how often a human removes what the scanner auto-added — the false-positive discovery loop, made systematic instead of accidental. Counts below #552's suppression threshold merge into one bare row.
| Term | Allergen | Auto-added | Verified | Removed | Level changed | Rejection rate | Businesses |
|---|
An allergen added by hand, with no scanner match at all — "which allergen are we blind to."
| Allergen | Missed count | Businesses |
|---|
| Allergen | Term | Kind | Evidence | Notes | Updated |
|---|
#556 (#551 M4) — across tenants, which Health-check findings are actually common, how often they're accepted rather than fixed, and how quickly they clear once they appear. Counts only, from each business's own weekly trend snapshot (#238) — no per-business detail here; that stays on the existing Health check screen and #500 parse-stats surfaces. This mechanism flows into product change (guidance copy, check calibration), never into runtime data.
| Check | Businesses affected | Findings outstanding | Accepted | Median days to clear |
|---|
#1985 (epic #1982) — how often one purchased thing is stored twice, and how much the same supplier product is called different things by different tills. Counts only, no receipt wording and no catalogue names: the duplicate figure is the one each business's own Health check computed (the single #785 matcher, read back out of its weekly snapshot), never recomputed here, and receipt names are counted inside the edge function and dropped. A blank duplicate count is not a zero, and the row says which kind of blank it is: no snapshot yet, usage stats off, or a snapshot that carries no count for this check. That last one is the honest answer to a real ambiguity, because a business with no duplicates and a snapshot written before this check existed look identical in the data: the client records a check only when it finds something.
| Business | Possible duplicate groups | Marked “not the same” | Receipt names learned | Products with 2+ receipt names | Worst |
|---|
Free text (#477) — no CHECK constraint. Distinct values with counts, so promoting a common one into the curated picklist (BUSINESS_TYPES in the app) is an informed, deliberate code change.
| Value | Businesses |
|---|
#2669 — each beta or trial tester (user and business), their in-app feedback, and sign-ins in the last 7 and 28 days. Silent signed in and raised nothing. Never signed in has no session. Admin-only via admin_silent_testers().
| Flag | Tester | 7d | 28d | Feedback |
|---|
| Time | Area | Message |
|---|