The Agent Passed the Test, but Broke the Workflow
About this session
A coding agent can make the test suite pass and still leave the engineering workflow worse than before. The code compiles, the pull request looks clean, and the local task appears “done.” Then CI fails in another service, a generated change violates an ownership boundary, observability disappears, or nobody can explain why the agent made the decision it made. This session is about the gap between agent success and workflow success. We will examine what breaks when coding agents move from individual tasks into real engineering systems: incomplete repository context, tasks that are too large for useful review, generated refactors that erase operational intent, pull requests that are hard to trust, and changes that pass local validation but fail under distributed system behavior. One concrete failure case covered in the talk is an agent-generated refactor that passed local tests but changed an operational contract outside the agent’s visible task boundary. The code-level task looked successful, but the workflow-level signals told a different story: dependent CI checks failed, review required rework, and observability signals needed for deployment confidence were missing or incomplete. The workflow-level metric that exposed the issue was not “agent task completion,” but whether the change could move safely through the engineering system: full-pipeline CI pass rate across dependent services, review rework rate, deployment validation success, and telemetry completeness after the change. Attendees will learn practical patterns for scaling agentic coding safely: context contracts, smaller task boundaries, CI and quality gates, ownership checks, reviewable diffs, bounded autonomy, action tracing, rollback-friendly changes, and evaluation criteria tied to engineering outcomes rather than demo completion.
Speaker
Key takeaways
- How to tell the difference between “the agent completed the task” and “the workflow is safe to merge, deploy, and operate.”
- How to structure coding-agent work with clearer context, smaller task boundaries, ownership checks, and quality gates.
- How to evaluate agentic coding workflows using reviewability, reliability, rollback safety, and production impact instead of local test success alone.
Related sessions
- AI Systems Are Missing a Trust Layer: How to Build Reliable AI in Production
- Squeezing the Token: A Lean Framework for LLM Cost and Prompt Optimization
- From Raw Telemetry to Automated Action: Building AI Pipelines That Auto-Resolve EV Charger Faults
- When RAG Meets Reality: Scaling Retrieval for Production