what happened?
pull the last 7 days into one structured state: funnel steps, failures, bounce, and feature usage.
field note · 19 sep 2026
i used one week of dupebrew product evidence to test jev as a decision layer. not a chatbot. not a session summariser. a small model that turns messy signals into choices, scores, and probabilities.
product analytics is good at telling me what happened. it can show the funnel, failure events, bounce, and which screens people reached.
but a list of signals is not a priority. someone still has to decide what to fix first, how urgent it is, and whether the evidence is strong enough to trust.
pull the last 7 days into one structured state: funnel steps, failures, bounce, and feature usage.
ask typed questions about priority, severity, the biggest drop, and evidence strength.
that split is the whole idea. jev never “reads sessions.” it receives a small state object and makes bounded decisions on top of it.
the dupebrew snapshot showed a familiar early-product problem: people reached the paywall, far fewer selected a plan, and the sample was still small.
there were also entitlement-sync and anonymous-user creation failures. those mattered, but the raw events needed interpretation before they became a product plan.
| question | jev type | why this shape |
|---|---|---|
| what should we prioritise first? | choice | one ranked focus, not another list |
| how urgent are the failures? | score | a bounded low → critical scale |
| where is the biggest drop? | choice | compare known funnel steps |
| is this sample strong enough? | score | keep confidence attached to the recommendation |
the important part is the boundary. jev chooses from options i define. it does not invent a roadmap in prose and make me reverse-engineer the answer.
fix the path where a purchase or restore finishes but premium access does not appear correctly.
make the plan choice and primary action easier to understand.
use the ranking for triage. keep collecting evidence before calling it a universal result.
jev returns confidence. the product decides when to act automatically and when a low-confidence result goes to review.
| signal | product action |
|---|---|
| entitlement reliability ranked first | make purchase / restore state sync the first fix |
| plan selection is the largest cliff | simplify pricing, default choice, and cta hierarchy |
| anonymous creation can fail | retry with backoff and avoid duplicate error noise |
| evidence is weak | treat the list as triage, then re-run it as the sample grows |
structured evidence in. typed questions out.
const result = await evaluate({
model: 'typesafe-ai/jev',
state: { /* PostHog funnel + error summary */ },
questions: {
primary_focus: {
type: 'choice',
instructions: 'What should we prioritize first?',
criteria: {
paywall_reliability: '…',
paywall_conversion: '…',
web_bounce: '…',
},
},
severity: {
type: 'score',
instructions: 'How urgent are purchase / entitlement failures?',
criteria: ['low', 'medium', 'high', 'critical'],
},
},
});the state contains aggregate product evidence, not raw session recordings.
jev is cheap and fast, but that is not the most interesting part to me. the useful part is that the output already looks like something a product process can use.
posthog says what happened. jev helps decide what to do next under uncertainty. that is a much better job for it than pretending it is another chat model.