September 25, 2026
In an era where generative AI promises to automate the entirety of the software development lifecycle, the role of the human engineer is undergoing a profound transformation. As algorithms increasingly write code, the burden of verification shifts from syntax to systemic integrity. In the latest episode of the Stack Overflow Podcast, host Ryan sits down with David Burns, the Head of Developer Advocacy and Open Source at BrowserStack, to dissect the evolving architecture of quality assurance.
The conversation moves beyond the hype of AI-driven coding to address a foundational truth: code is only as reliable as the tests that validate it. For Burns, a veteran of the automation space, the solution to modern software fragility lies in a return to first principles—professional skepticism, disciplined testing frameworks, and the rigorous management of state.
The Core Thesis: Professional Skepticism in the Age of AI
As Large Language Models (LLMs) become integrated into IDEs, developers are shipping features at unprecedented speeds. However, this velocity introduces a "black box" problem. When an AI generates a complex function or an entire microservice, the human developer often assumes the code is functional because it appears syntactically correct.
David Burns argues that this is where the industry faces its greatest risk. "We are entering an age where the ability to generate code has outpaced our ability to verify it," Burns notes. He advocates for "professional skepticism"—a mindset where developers treat AI-generated output as a draft that must undergo the same, if not more rigorous, scrutiny as code written by a junior developer.
This skepticism is not merely about debugging; it is about architectural validation. Burns posits that if an AI cannot explain the "why" behind its logic, the engineer must treat that code as a liability. In an automated world, the developer’s primary value proposition is shifting from writing code to auditing the logic that drives systems.
Chronology: The Evolution of Testing Paradigms
To understand where we are going, we must look at how the landscape of software quality has shifted over the last two decades.
- The Manual Era (Pre-2010): Quality assurance was largely a human-centric endeavor. QA engineers manually clicked through interfaces, a process that was slow, error-prone, and unsustainable for the rapid release cycles of the early web.
- The Rise of Automation (2010–2020): Tools like Selenium revolutionized the industry. Developers began writing scripts to simulate user interactions, moving testing "left" in the development cycle. This era birthed the "BrowserStack" approach—testing across a fragmented landscape of browsers and devices.
- The Agentic Turn (2020–2026): We have moved into the era of agentic engineering. Autonomous agents are now capable of navigating applications, identifying elements, and executing workflows. However, as Burns points out, these agents are often brittle because they lack an understanding of the underlying application state.
Supporting Data: Why Flaky Tests Persist
A recurring theme in the discussion is the "flaky test"—a test that passes or fails inconsistently despite no changes to the code. According to industry surveys, flaky tests are the single biggest drain on developer productivity in 2026.
Burns contends that flakiness is rarely a tool issue; it is a state management issue.
The Anatomy of a Flaky Test
- Asynchronous Dependencies: Modern applications rely on distributed microservices. If a test suite doesn’t account for the latency of these dependencies, it will fail intermittently.
- Shared Global State: When tests modify database records or global variables without proper teardown procedures, subsequent tests inherit a corrupted environment.
- The "Race Condition" Fallacy: Developers often attempt to "fix" flakiness by adding arbitrary wait times (
sleepcommands). Burns describes this as a "code smell" that hides underlying architectural flaws rather than resolving them.
By applying Test-Driven Development (TDD) to agentic workflows, developers can force agents to operate within strict state boundaries. Burns suggests that we should treat agentic operations as deterministic processes, where every action is followed by a state-validation check.
Official Perspectives: The BrowserStack Vision
BrowserStack has positioned itself at the epicenter of this testing evolution. As the Head of Developer Advocacy, Burns is not just observing these trends; he is building the infrastructure to support them.
"Our goal at BrowserStack is to remove the friction between development and verification," Burns explains. The company’s focus has shifted from simple browser compatibility to providing an observability layer for automated agents. By giving developers visibility into what an agent sees and how it processes the DOM (Document Object Model), BrowserStack is enabling a new class of "intelligent testing."
Burns emphasizes that the open-source community remains the heartbeat of this movement. Through BrowserStack’s Open Source Office, the company continues to sponsor projects that maintain the standards for cross-browser and cross-platform automation. He argues that if we do not standardize how agents interact with the web, we will face a future of fragmented and unmaintainable test suites.
Implications: The Future of Agentic Engineering
What does this mean for the average developer? The implications are three-fold:
1. The Rise of the "Quality Engineer"
The distinction between developer and QA will continue to blur. Future engineers will be expected to be masters of both production code and the testing harnesses that govern it. Proficiency in writing "testable" code will become a high-demand skill, potentially more valuable than the ability to write raw feature code.
2. Deterministic AI
The industry is moving toward "Agentic TDD." In this paradigm, an AI agent is instructed to write a test before it writes the feature. This forces the agent to define the boundaries of the system and the expected state changes, effectively creating a "guardrail" for the AI’s generative capabilities.
3. State Management as the New Frontier
As applications become more complex, the ability to manage state—at the database level, the UI level, and the API level—will be the primary differentiator between reliable systems and "flaky" ones. Developers who master distributed state management will find themselves at a distinct advantage.
Community Recognition: Celebrating Excellence
In the spirit of fostering a healthy ecosystem, the conversation concluded by recognizing individual contributions to the developer community. Stack Overflow recently awarded the "Populist" badge to user brentvatne for their insightful solution regarding the programmatic reading of app.json and exp.json files.
This recognition serves as a microcosm of the larger theme discussed by Burns: the importance of sharing knowledge and solving the "niche" problems that prevent software from working as intended. Whether it is a configuration issue in a React Native app or a complex integration test, the collective intelligence of the developer community remains the most robust tool we have against the uncertainty of the AI age.
Conclusion: A Call to Rigor
The shift toward AI-driven development is inevitable, but its success depends on the discipline of the humans who oversee it. David Burns’s message is clear: do not let the speed of AI lull you into a false sense of security. Embrace professional skepticism. Invest in your testing frameworks. And above all, master your application’s state.
As we look toward the end of 2026, the developers who thrive will not be those who rely solely on AI to write their code, but those who use AI to accelerate their testing and validation processes. Quality is not an afterthought; it is the foundation upon which all modern software is built.
For those interested in diving deeper into these topics, David Burns invites you to follow his insights on LinkedIn and subscribe to the BrowserStack Talks podcast.








