What stakeholder demos caught that our tests couldn't
Four kinds of failure our test suite couldn't see, which only surfaced once real stakeholders used an AI booking assistant in a room.
Every older note, newest first.
23 notes in all
Four kinds of failure our test suite couldn't see, which only surfaced once real stakeholders used an AI booking assistant in a room.
A tooling migration, some dependency bumps and a telemetry schema: none of it on a roadmap, and all of it about keeping several frontend services easy to work in.
Building a standardised logging system for an AI assistant service: making consistent structured logging the path of least resistance, and thinking carefully about what an AI service should and shouldn't persist.
Merged pull requests already record what you did. Turning them into a work log with an LLM is easy; the hard part is the defensive rules that stop it corrupting the vault.
Removing unused documentation tooling and fixing prop type warnings: small changes that don't ship features but keep the console's warnings meaningful.
A currency display bug, a feature flag that never existed in the flagging service, and a passenger type code nobody resolved: three kinds of drift that fail silently.