Shipping Software While You Sleep: Inside an Overnight Agentic Dev Loop

About this session

Everyone's asking whether AI can write production code. The more useful question is what has to be true around the model for agent-written code to actually ship — and that's a pipeline problem, not a prompt problem. This session walks through an agentic development pipeline built by a small engineering team to take work from a rough ticket to tested, reviewable code with minimal human routing. We'll cover the parts that make it real: structured ticket decomposition and readiness scoring so agents get unambiguous specs, an overnight coding loop that does the first pass while the team is offline, and — most importantly — the testing and evaluation architecture that decides whether agent output is trustworthy enough to merge. The honest core of the talk is the guardrails: ephemeral isolated environments, automated end-to-end testing, and eval-based quality gates, because none of this works without them. You'll also hear what changed about engineering roles, where the pipeline saves real time, and where it quietly creates new failure modes you have to design against.

Speaker

Key takeaways

  • Key takeaways: A concrete pattern for turning vague tickets into agent-ready specs (decomposition + readiness scoring) — the input quality problem most teams skip. How to make agent-written code shippable: ephemeral environments, automated E2E testing, and eval gates as the trust layer. What an agentic pipeline actually changes about team structure, throughput, and the new risks it introduces.

Related sessions