The Big Fat Geek

Personal blog of Prasad Ajinkya

AI in QA: Maturity Curves, the Shared Brain, and the OpenTelemetry Blindspot

AI in QA Shared Brain Architecture

Every technology conference agenda features them prominently: “The Future of Enterprise Artificial Intelligence,” “Navigating Disruption in 2026,” or “Scale, Strategy, and Digital Transformation.” You balance a lukewarm coffee, locate a seat in the auditorium, and confront the familiar dilemma: Is this going to be an unvarnished masterclass in production engineering, or forty-five minutes of synchronized corporate throat-clearing?

Public panel discussions represent the staple of industry gatherings, yet they remain notoriously polarizing. At their worst, they descend into polite theater—four panelists taking turns reciting safe public-relations platitudes while an event sponsor reads canned prompts off a tablet. But when you move from public auditoriums into invite-only, closed-door executive roundtables, the dynamic inverts completely.

Recently, I participated in an off-the-record boardroom roundtable on Artificial Intelligence in Quality Engineering and Assurance (QA) alongside engineering leadership from premier financial institutions. When commercial posturing is stripped away, you gain a rare, unvarnished look at where enterprise adoption actually stands, stress-test your architectures against peer implementations, and confront profound industry-wide blindspots.

The Four-Stage AI-in-QA Maturity Curve

In institutional banking and regulated finance, where audit trails and statutory invariants are mandatory, non-deterministic models cannot simply be unleashed unsupervised into production release trains. The roundtable revealed a distinct evolutionary progression across leading banks:

The Four-Stage AI-in-QA Maturity Curve: Tier-1 institutions progress through four distinct phases: Level 1: Syntactic boilerplate test generation → Level 2: Semantic test synthesis from PRDs → Level 3: Context-aware blast-radius testing → Level 4: Closed-loop observability with automated root-cause healing.

  • Level 1: Syntactic Code Helpers (The Baseline): Engineering teams deploy AI copilots to generate boilerplate unit tests and translate manual test scripts into Playwright or Cypress suites. Every bank has completed this step, but participants conceded that generating more code merely balloons technical debt when user interfaces shift.
  • Level 2: Semantic Test Synthesis: Feeding Product Requirement Documents (PRDs), user journeys, and OpenAPI schemas into models to synthesize multi-variable edge cases automatically.
  • Level 3: Context-Aware Autonomous Testing (The Frontier): Testing engines that comprehend distributed microservice dependencies, calculate the blast radius of git pull requests, and prioritize test execution based on historical defect clustering.
  • Level 4: Self-Healing & Closed-Loop Observability: Autonomous harnesses that ingest live production telemetry, detect regression signatures, construct synthetic reproduction tests, and propose patch PRs prior to human triage.
“Traditional automated QA relies on black-box assertions. But when autonomous AI agents plan test journeys, black-box testing breaks down. Without OpenTelemetry, you are searching for needles in haystacks.”

Architectural Validation: The Shared Brain for QA Agents

A striking takeaway from the session was hearing leaders at the frontier describe an architectural paradigm I have championed across our systems at Homeville: the Shared Brain and deterministic memory.

The primary point of failure in early AI testing is stateless isolation. If an autonomous test agent is dispatched to validate an escrow reconciliation microservice without historical context, it is blind to critical invariants:

  • How statutory escrow settlement rules intersect with Reserve Bank of India co-lending directives.
  • Which downstream banking partner APIs produce intermittent socket timeouts on the first banking day of the month.
  • The architectural invariants recorded in historical Architecture Decision Records (ADRs).

The engineering organizations securing genuine production return on investment are not relying on ad-hoc prompts. They are curating a centralized context repository—a shared organizational memory unifying service schemas, post-mortem incident logs, regulatory mandates, and dependency graphs. Autonomous QA agents query this repository prior to test generation, operating with the accumulated institutional wisdom of the entire department.

The Industry Blindspot: OpenTelemetry (OTel)

While the boardroom was animated regarding model benchmarks and synthetic data harnesses, a pivotal architectural query was met with widespread silence:

“How are you instrumenting your autonomous AI test agents with OpenTelemetry (OTel) to trace reasoning spans and correlate agent tool failures directly with backend distributed traces?”1

Virtually no institution in the room had architected for this. Traditional automated testing operates on simple black-box assertions: dispatch request, assert HTTP 200, verify DOM element. But when an autonomous agent is evaluating journeys, calling external APIs, and dynamically altering inputs, black-box assertions are completely inadequate.

1. Isolating Non-Deterministic Agent Failures

When an autonomous test agent fails to complete a journey, where did the breakdown occur? Was it a genuine software defect in the application, a hallucinated tool argument during the model’s planning phase, a rate-limited embedding API, or a silent database lock in a downstream microservice? Without OpenTelemetry trace spans linking the agent’s internal reasoning steps directly to backend service traces, debugging agent failures is an exercise in futility.

2. Bi-Directional Telemetry: Feeding Production into QA

OpenTelemetry must not remain confined to production monitoring; it should directly power quality assurance. Real-world user paths, payload distributions, and latency anomalies captured in production traces should automatically feed the AI testing engine, enabling it to synthesize hyper-realistic regression scenarios rooted in actual user behavior.

3. Token Economics and Prompt Drift

Executing thousands of agentic test journeys across continuous integration pipelines incurs non-trivial token expenditures. Tracking context bloat, prompt drift, and model inference latency across releases requires the identical semantic conventions and telemetry standards we apply to distributed microservices.2

Strategic Takeaways: The Value of the Right Room

This dynamic returns us to the utility of peer executive roundtables. Public panel discussions offer marketing narratives and corporate safe harbors. But closed-door assemblies with authentic practitioners provide the engineering reality.

They validate foundational architectural choices—such as anchoring agent intelligence in a version-controlled Shared Brain—while exposing massive, systemic blindspots like OpenTelemetry before they harden into crippling technical debt.


  1. OpenTelemetry (OTel), an open-source observability framework under the Cloud Native Computing Foundation (CNCF), providing vendor-neutral APIs, SDKs, and tooling to generate and export telemetry data (traces, metrics, logs).
  2. OpenTelemetry Semantic Conventions for Generative AI define standardized telemetry attributes for tracking prompt token usage, completion latencies, model provider identities, and tool invocation spans.