Automating MCP Server Testing: Engineering Reliability for Agentic Systems

About this session

AI agents don’t fail like traditional software. They don’t just throw exceptions, they drift, misinterpret tools, invoke the wrong functions or behave differently across environments. When deploying Arm’s Open-Source custom MCP server to power AI assistants for architecture development, migration, & optimization, we faced a critical question: how do we test a system built for nondeterministic interaction? In this talk, I’ll share how we moved from manual validation to a repeatable, CI-enforced testing strategy using Pytest & Testcontainers. We spin up real MCP server in Docker during tests, validating tool discovery, invocation, & protocol compliance end-to-end. This isn’t about mocking LLM output. It’s about testing the contract between agents & tools. The key insight: treat your MCP server like production infrastructure, not experimental glue code. Because “it worked on my machine” is not a deployment strategy. This session contributes practical, open, & reproducible testing approach that any project can adopt to improve reliability & trust in Open-Source MCP implementations. Thus, making agent infrastructure production-ready.

Speaker

Key takeaways

  • Recognize why unit tests are insufficient for agent-facing systems
  • Learn how to run MCP servers inside containerized test environments through demonstration of Arm’s Open-Source MCP server(github.com/arm/mcp)
  • Learn how GitHub Actions can be used to automates CI integration testing

Related sessions