Writing

Notes on product, AI and automation

I write about product culture, AI-assisted development, automation and measurement, drawing on real experience, mistakes and systems I have built. The throughline is my shift from Product Manager to AI Product Builder: what changes when the person deciding what to build can also build it, while staying in control of the process and the decisions.


Product culture and Fractional PM

MVPs need atomic-unit tests

A practical scope ritual for founders who need a smaller MVP: name the core object before adding feature bundles.

Product discovery needs underserved-need ladders

Turn discovery notes into a ranked ladder of segment, need, gap, evidence, and next test before roadmap choices harden.

Product teams need asymptote registers

When growth flattens after early fit, stop polishing the old curve. Use an asymptote register to test the next ceiling.

OKRs need heartbeat metrics

Do not force every team into fake OKRs. Separate strategic change from heartbeat work with health metrics and escalation rules.

Category strategy needs behavior maps

Positioning a new category starts by mapping the behavior customers must adopt, not by polishing feature claims.

Planning needs public evidence, not confident forecasts

A practical ritual for product teams: record the claim, seek outside evidence, update confidence, and decide before the roadmap hardens.

Product bets need kill criteria, not optimism

Turn roadmap bets into decision contracts with stop, change, and continue criteria before optimism becomes inertia.

Product strategy needs choice maps, not priority stacks

A practical way to connect product strategy to roadmap bets by mapping choices, exclusions, capabilities, and review triggers.

Roadmaps need bet ledgers

Replace feature-calendar debates with a roadmap ledger that makes evidence, ownership, confidence, signals and decisions explicit.

Product operating models start with decision rights

A practical way to turn product operating model talk into clear ownership, evidence, risk approval, and follow-through.

Discovery interviews need decision logs

Customer interviews only matter when they change product decisions. Use a lightweight log to connect evidence, assumptions, and ownership.

Your trial is not broken: you are bringing in the wrong users

Before rebuilding onboarding and nudges, a product team needs to know whether signups are actually in the target. Often the leak starts before the trial.

Product Operating Model: roles are not enough

A Product Operating Model is not installed through new job titles. It changes how teams decide, learn, measure and own outcomes.

Product prioritization: roadmaps cannot hold everything

Prioritization does not come from a perfect score. It comes from explicit trade-offs between outcomes, risk, evidence and team capacity.

Discovery is not asking customers what they want

Discovery reduces uncertainty before a team builds. Asking customers what they want and turning it into roadmap is not enough.

What is an asset, really, for a company?

A tool can help you work faster, but an asset compounds over time. The difference matters when teams decide what to protect, automate and build.

AI Product Builder

AI roadmaps need task-exposure maps

Before adding AI to a workflow, map exposed tasks, human input, failure consequences, and the right rollout choice.

AI coding harnesses need blast-radius audits

Before trusting a coding agent in a real repo, audit the harness that can read, edit, execute, install, update, and transmit.

Agent releases need trajectory gates

Before shipping an agentic feature, gate the expected tool path, tolerated variance, groundedness, safety, and regressions.

AI-assisted coding needs ambiguity checkpoints

AI-generated code is often mostly right. Teams need a checkpoint that captures ambiguity before review and ownership.

AI evals need fail-case budgets

A practical release gate for LLM products: replace green score dashboards with real fail cases, aligned evaluators, and reruns.

AI product milestones need demo-to-product gates

A practical launch gate for turning AI demos into reliable workflows without mistaking early success for product readiness.

Model selection needs task portfolios

Choosing an AI model starts with a task portfolio: quality, latency, cost, fallback, owner and review thresholds.

AI coding agents need review packets

Safer AI-generated code starts with a review packet: intent, touched surfaces, tests, failures, owner, and rollback path.

AI coding agents need rollback plans

Reviews matter, but production safety needs ownership, blast-radius control, stop conditions, and a rollback path.

AI coding needs acceptance criteria first

AI coding agents become easier to trust when every task starts with a goal, constraints, acceptance criteria, review path, and rollback note.

AI confidence theater: build workflows that hold

The problem is not using too little AI. It is confusing demos, agents and token burn with workflows that actually change delivery.

Fable 5 is back: benchmarks, stops and safeguards

The Fable 5 case shows that benchmarks matter only when we connect them to access, safeguards, evals and real team workflows.

After vibe coding: model routing makes AI coding governable

Model routing moves AI coding from individual enthusiasm to a governed system: tasks, costs, fallbacks and review become product decisions.

Inference, context, evals: AI principles before models

Choosing the right model matters, but a useful AI product depends on inference, context, harnessing, evals, fallback and review.

An exobrain is only as good as the context you give it

Everyone has access to the same AI now. The edge is not the exobrain itself: it is the quality of the strategy you put inside it.

AI-assisted delivery: more output does not mean more control

Coding agents can make delivery faster. Without clear review criteria, ownership and stop gates, they can also increase the noise.

AI automation systems

Agent memory needs dose tests

Before expanding persistent memory, test how each agent handles baseline, retrieval, full context, token overhead, and saturation.

Local agents need context budgets

Before raising a local agent context setting, write a budget for inputs, traces, answer space, memory, and fallback behavior.

AI workflows need escalation lanes

Stop making humans babysit every AI step. Route work by risk, approval need, timeout, and evidence trail instead.

Production agents need runtime guardrails

LLM agents in production need spend caps, rate limits, fallbacks, redaction, and trace review as one runtime contract.

RAG agents need evidence handoffs

Make agentic RAG auditable by turning every retrieval step into a claim-level evidence handoff, not hidden plumbing.

Structured outputs need contracts, not parser patches

LLM workflows become reliable when teams define output schemas, validation, fallback paths and ownership before adding retries.

AI agent observability needs trace contracts

Screenshots do not explain agent behavior. Production AI agents need trace contracts that show actions, tools, approvals, and owners.

AI agents need tool permissions, not blanket access

Internal AI agents become safer when every tool action has scoped data access, approval rules, logging, and an owner.

RAG needs retrieval contracts, not bigger windows

Reliable RAG workflows come from a retrieval contract: source scope, ownership, freshness, filters, ranking, fallbacks, and evals.

AI automation needs failure classes, not retries

Retries are useful, but production AI workflows need named failure classes, owners, and routing rules before they can scale safely.

Retrieval needs a contract, not just a vector database

Reliable RAG starts before the answer: teams need a retrieval contract for sources, freshness, fallbacks, ownership, evaluation, and review.

Agents need queues, not just prompts

As agents move into long-running work, the product problem shifts from prompt quality to queues, ownership, WIP limits and review.

Automation means turning signals into actions

An AI stack creates no value if signals stay still. The real work is connecting data, triggers, ownership and human review.

LLM "hallucinations" are often retrieval failures

Many LLM hallucinations start before generation: weak retrieval, missing context, and no rule for what the model should do when evidence is thin.

RAG or CAG: the operational choice before the architecture

RAG and CAG are not interchangeable labels. The choice depends on knowledge size, freshness, latency and governance.

n8n in production: the workflow is not done when it runs

A useful n8n workflow has ownership, logging, error handling and fallback. The value is not in the demo, but in daily operations.

Measurement and operational autonomy

Adobe Mix Modeler vs Meridian vs Robyn: choose by readiness

A practical comparison of Adobe Mix Modeler, Meridian and Robyn for teams choosing a measurement operating model, not just a package.

Digital metrics need boundary registers

Before building a digital transformation dashboard, define what each metric includes, excludes, proxies, delays and decides.

AI risk reviews need evidence receipts

Maturity scores do not prove that AI controls work. Use evidence receipts to connect risks, owners, thresholds and decisions.

AI adoption metrics need denominator maps

A useful AI adoption dashboard maps eligible, tried, active, blocked, abandoned, embedded, and governed use.

Measurement systems need score receipts

Automated scores are decision artifacts. Give each consequential score a receipt that explains inputs, thresholds, owners, and appeal paths.

Growth metrics need survival cohorts

A growth dashboard should track births, deaths, reactivations and survival cohorts, not only the newest accounts.

Anonymized datasets need evidence tests

Before analytics data enters AI workflows, replace the anonymized label with a testable evidence record for residual risk.

AI systems need security ledgers, not checklists

For AI features using personal data, security must be a living ledger across data, models, owners, evidence, and shutdown paths.

Consent mode needs fallback states, not panic

Keep GA4 useful after consent changes by defining fallback states for data, confidence, and decision ownership.

Experiment readouts need decision rules

Experiment reviews should not reopen metric politics. Pre-agreed decision rules turn A/B test readouts into accountable action.

MMM need decision boundaries, not attribution theatre

Use marketing mix modeling as decision support by defining what the model can decide, when to wait, and who owns budget action.

Measurement Protocol is not a tracking plan

GA4 Measurement Protocol improves event ingestion, but it cannot replace the discipline of defining meaning, ownership, validation, and decision use.

Govern AI agents by autonomy level, not by vendor or model

The same model can be harmless or risky depending on the autonomy you give it. The real question is access, blast radius and approval level.

MMM with Meridian and Robyn: data comes before the model

Meridian and Robyn make Marketing Mix Modeling more accessible, but the real work is still in data, assumptions and causal interpretation.

GA4 Measurement Protocol does not replace tagging

Measurement Protocol enriches GA4 with server-side and offline events. Treating it as a tagging replacement creates weak data.

Attribution is a measurement contract

Attribution is useful when teams know the rules, limits and purpose of the model. It becomes risky when treated as objective reality.

Measuring outcomes when the team still looks at hours

If the default metric is still time spent, AI and automation can increase activity without improving outcomes.

An operational dashboard is an internal product

The best operational dashboards help teams act, not just read numbers. They should be designed like internal products.