Context
I worked on a checkout and payments platform where the same operational questions came back every week: is this payment method failing for merchant X, did that order really fail, is this behaviour caused by a feature flag. The answers existed, but they were spread across tools with different auth models and different query languages: observability (Sentry, Grafana, Prometheus, Loki), product analytics (Mixpanel), feature flags (GrowthBook), issue tracking (Linear), Slack and a set of operational runbooks.
Engineers were spending part of their day being a search engine across six dashboards.
Problem
- Low-complexity operational questions consumed engineering time meant for product work.
- Answering them required knowing which tool to look at and how to query it.
- Several sources hold personal data, so the automation had to be safe by construction, not safe by policy.
- Answers had to be defensible: incident response is only as good as the evidence behind it.
My responsibility
I developed ADA: the agent service, the tool integrations, and the permission model that decides what the agent is allowed to read.
What I implemented
- An agent service in TypeScript/Node.js, deployed on Vercel, using LLMs through Amazon Bedrock (Claude).
- Tool integrations with Slack, Sentry, Grafana, Prometheus, Loki, Mixpanel, Linear and GrowthBook, a mix of custom tools and tools exposed through MCP.
- A read-only access model: the agent may inspect state, never mutate it.
- Authentication with OIDC/OAuth plus granular per-tool permissions, so the agent's reach is explicit rather than implicit.
- Personal-data protection on what each tool is allowed to return.
- Feature flag analysis, turning "is this flag responsible?" into a question answered with data.
- Automated incident investigation: correlating signals across tools before a human starts digging.
- A knowledge base of operational runbooks the agent consults instead of improvising procedure.
- A daily hypercare digest per merchant, delivered on a schedule so issues surface before someone opens a ticket.
How it works
A request arrives from Slack. The agent plans which tools to call, calls them through the authenticated tool layer, and composes an answer from the returned evidence. Because the tools are read-only and scoped, the worst consequence of a bad tool choice is a useless query, not a production change. The same service is reused by a scheduled job that produces the daily per-merchant digest.
Design decisions and trade-offs
- Read-only, on purpose. The agent cannot act on production. Trade-off: it cannot remediate automatically, so a human still executes the fix. I accepted a slower loop in exchange for an automation that can never cause an incident.
- Integrate the tools teams already use instead of building a new store for observability data. Trade-off: we inherited each vendor's rate limits, query language and auth quirks, and the agent's freshness depends on those tools.
- Granular permissions instead of one broad credential. Trade-off: more configuration and more moving parts, but a misused or leaked tool cannot read everything.
- Runbooks as explicit knowledge rather than trusting the model to recall procedure. Trade-off: runbooks must be maintained, or the agent will confidently follow a stale process.
Result
Roughly 70% fewer support tickets reaching the development team after ADA started handling operational questions and first-pass incident triage.
What I learned
- In agent design the constrained surface (what a tool is allowed to do) matters more than the prompt.
- Read-only and least privilege are engineering decisions, not policies: they have to be encoded in the integration layer.
- An agent is only useful when its tools return real evidence; otherwise it produces plausible text that nobody can act on.