Automating MCP Server Testing: Engineering Reliability for Agentic Systems
About this session
AI agents don’t fail like traditional software. They don’t just throw exceptions, they drift, misinterpret tools, invoke the wrong functions or behave differently across environments. When deploying Arm’s Open-Source custom MCP server to power AI assistants for architecture development, migration, & optimization, we faced a critical question: how do we test a system built for nondeterministic interaction? In this talk, I’ll share how we moved from manual validation to a repeatable, CI-enforced testing strategy using Pytest & Testcontainers. We spin up real MCP server in Docker during tests, validating tool discovery, invocation, & protocol compliance end-to-end. This isn’t about mocking LLM output. It’s about testing the contract between agents & tools. The key insight: treat your MCP server like production infrastructure, not experimental glue code. Because “it worked on my machine” is not a deployment strategy. This session contributes practical, open, & reproducible testing approach that any project can adopt to improve reliability & trust in Open-Source MCP implementations. Thus, making agent infrastructure production-ready.
Speaker
Key takeaways
- Recognize why unit tests are insufficient for agent-facing systems
- Learn how to run MCP servers inside containerized test environments through demonstration of Arm’s Open-Source MCP server(github.com/arm/mcp)
- Learn how GitHub Actions can be used to automates CI integration testing
Related sessions
- Nasdaq’s Agentic Journey: A Case Study in MCP-Driven Innovation
- Your Inbox Is the New Interface: Building AI Agents on Chat applications
- Agents Don't Fail in the Sandbox: What Actually Breaks When You Ship AI Agents to 1000's of Users
- AI-GENERATED VOICE SYSTEMS: INTEGRATING COPYRIGHT LICENSING, BIOMETRIC DATA PROTECTION AND AUTOMATED ENFORCEMENT MECHANISMS