Connect with us

AI

The AI Boom Still Needs Someone Watching the Agents

PagerDuty’s Jennifer Tejada says tiny AI teams will still need a third party watching their agents, because drift fails quietly while the SaaS seat model cracks.

Published

on

Jennifer Tejada says AI agents will not fail like crashed servers, and that is why tiny production teams will still need an independent watcher. The former PagerDuty chief, now executive chair, has lived through the late-1990s internet buildout and the cloud years, and she is treating the move from AI experiments into live systems as the next version of that pattern, with a quieter break in it.

The boom she is celebrating is the same force that makes the old red-alert model late. Smaller teams can ship more software. Those teams also run agents whose mistakes look like normal work until the damage has already spread.

Tejada Handed Over the CEO Job and Kept the Warning

PagerDuty named John DiLullo chief executive on May 11, 2026, and moved Tejada to executive chair after she had run the company since 2016. She took it public on April 11, 2019, and is still on the board of The Estée Lauder Companies. In San Francisco she has been watching AI listings land in a city that already went through one internet gold rush and one cloud cycle.

DiLullo, who previously ran Deepwatch, LiveVox, and Lastline, said AI-assisted software work and more tangled digital estates will expand the need for real-time operations management. That is the vendor’s own bet, and it matches the warning Tejada kept repeating after she gave up the day-to-day job: someone has to watch the agents, and it should not be the agents marking their own homework.

PAGERDUTY AFTER THE HANDOFF

  • Customers: The company says it serves more than 35,000 customers, including two-thirds of the Fortune 100, and pulls signals through over 700 integrations.
  • Tenure growth: Under Tejada it scaled revenue from under $50 million to nearly $500 million.
  • FY2027 outlook: On the same day as the CEO change, it restated full-year revenue guidance of $488.5 million to $496.5 million, with a first-quarter range of $118 million to $120 million.
  • The new brief: Tejada said she wants to spend more time on the AI boom while DiLullo runs the shop.

She still talks like an on-call executive. Nvidia is a customer, and she described walking into a meeting to introduce DiLullo and finding a class of interns already in the building. Her point was that the hiring bar for this wave is not a decade of ops scars. The other half of that thought is less friendly to the intern: the systems those new teams will touch fail in ways a junior pager rotation was never trained to see.

Fifteen Hours When a DNS Bug Went Global

When Tejada wants a picture of a loud failure, she reaches for Amazon Web Services in October 2025. That event was not an AI agent going rogue. It was automation inside DynamoDB’s DNS machinery in US-EAST-1, Amazon’s oldest and busiest region, hitting a race condition and wiping the regional endpoint’s IP addresses. Customer traffic and internal AWS services lost the path at the same time.

Amazon later said the cascade lasted 15 hours and 32 minutes. Ookla said its DownDetector service took more than 17 million reports of trouble at 3,500 organizations, with the heaviest volumes in the United States, the United Kingdom, and Germany. Snapchat, AWS itself, and Roblox sat at the top of that list. Even after DynamoDB came back, EC2 spent hours chewing through a backlog of network state, and load-balancer health checks failed on the way up.

THE AWS CLOCK IN OCTOBER

  1. October 19, 2025, 11:48 p.m. PDT: DynamoDB API error rates rise in US-EAST-1 after a delayed DNS Enactor and a second Enactor’s cleanup leave the regional endpoint with an empty record.
  2. October 20, 2025, early morning: Engineers restore DynamoDB by hand, then EC2’s lease machinery tips into congestive collapse and new instances launch without working network paths.
  3. October 20, 2025, afternoon: Network load-balancer health failures spread connection errors into Lambda, Fargate, Redshift, and the support console before Amazon calls the fleet back to normal.

That is the failure mode the industry already knows how to staff. Dashboards go red. Status pages fill. Humans get paged. Tejada’s claim is that this is the easy case. A classic outage stops. An agent that has drifted keeps working, and it can keep working wrong in several directions before anyone treats it as an incident.

An Agent Off the Rails Never Turns a Dashboard Red

Software that dies is obvious. Model output that slowly gets worse is not. Tejada called that gap the reason early detection has to sit beside the usual infrastructure monitors, and why AIOps vendors now talk about watching models with the same seriousness they once reserved for disks and queues.

When AI drifts, it’s actually harder to see, and you don’t see until it’s executed that drift in a number of ways, and now it’s evolved into multiple failures.

Jennifer Tejada, Executive Chair, PagerDuty

PagerDuty’s own engineers put the same idea in blunter shop language on July 21, 2026, weeks after she gave that interview. An AI agent that has quietly gone off the rails does not turn a dashboard red. Their monitors, built on LLM-as-a-judge scores, often mean “something might be broken,” not “the site is down.” A relevance drop is usually a prompt change, a retrieval tweak, or a model swap, not a rollback you can finish in fifteen minutes. It can also be an upstream outage dressed up as a quality slide, which is why they want one place that can split true output drift from a ordinary app failure.

The public argument about agents still prefers a movie. Multi-agent swarms, hidden prompts, collusion. Builders who actually ship this stuff keep pointing at a duller break: orchestration, state, retries, and a model that never throws an exception. The cheaper on-call agents now being pitched against Datadog and PagerDuty even claim classic tools wake people for faults that already fixed themselves. That complaint does not kill Tejada’s thesis. It confirms the next layer of the irony. The watcher is also becoming an agent, and someone will have to watch that too.

HOW THE SRE AGENT TAKES THE FIRST PASS

  • Classify: It decides whether an AI incident should wake a human at all.
  • Remediate: It can run pre-approved fixes without waiting for a person to click through a runbook.
  • Ticket: It files follow-up work for the team that owns the model or the prompt.
  • Trace: Tied into tools such as Arize, it pulls the LLM calls, tool calls, retrievals, and outputs behind a failing request and hands the responder a ranked next step.

Tejada’s product version of that loop is a third party with permission to interrupt. “Over time humans are going to want something watching their agents for them and giving them an unbiased opinion of how their agents are working,” she said, “and enabling them to disrupt or interrupt their agents’ work if they see something that they don’t like.” Unbiased, in that sentence, means not graded by the same system that just took the action.

What a One-Pizza Team Cannot See

Amazon used to describe a two-pizza team as a group small enough to feed with one or two pies. Tejada says AI will push companies toward one-pizza and two-pizza teams as the default. Fewer people. More software in production. More agents sitting on tools, tickets, and customer paths, including consumer agents that shop and check out.

A six-person team can ship an agent that talks to refunds, search, or a hospital inbox. It cannot, at 2 a.m., reconstruct why that agent’s tool-choice score fell, whether the eval itself is noisy, and whether the right move is a prompt edit or a model fallback. The old answer was to hire more people onto the pager. The new answer, in her telling, is to buy a watcher and keep the team small anyway.

That is the part of the AI IPO story that does not show up in ribbon cuttings. Headcount is supposed to fall. Complexity is supposed to rise. The gap between those two lines is where silent failures live. If the team is too small to notice drift, the company either accepts mystery regressions or it pays for a control plane that is willing to stop the agent.

She still wants the upside. Healthcare and transportation were the examples she reached for, and the hospital version of this problem is already in the building: hospitals already running AI without a map are a preview of pizza-team production without a watcher. Failure, she said, is also how you get resilience. The rule she would actually run is narrower than that slogan.

Seat Prices Cracked While the Control Plane Held

The same months that put agents into production also repriced the software those agents were supposed to replace. In early 2026, Jefferies tagged a slide in application software as the SaaSpocalypse, after a scare that agents would erase the human seats most SaaS vendors still bill on. Names in that group fell 30 percent to 55 percent. The panic treated every workflow product as a menu an agent could skip.

A lot of that tape later bounced, because agents still need records, permissions, audit trails, and a place to write the answer down. What did not bounce as easily is the unit of account. Seats look optional. Consumption, protection per workload, and operations control do not, at least not to the buyers who just watched a quality metric rot without a red light. PagerDuty is trying to live on that second list, next to identity, recovery, and endpoint tools that already sold themselves as infrastructure rather than as another login.

THREE PRESSURES ON THE SAME STACK

Pressure How you notice Who is on the hook
Classic cloud outage Status pages and pagers light up On-call humans and the cloud vendor
Agent drift Answers get worse and nothing crashes Whoever is paid to interrupt the agent
Seat-model shock Software stocks reprice, not servers App vendors still charging per login

Tejada is not a neutral witness here. She spent a decade selling the pager, and she is now selling the idea that the pager has to watch models too. The reason the pitch still lands is that the other two columns already happened. AWS did go dark for a working day because an internal automation race went badly. Software stocks did eat a seat scare. And PagerDuty’s own staff now write, in public, that a bad agent will not do them the courtesy of looking like an outage.

Who Still Has to Show Up at the Data Center

The other half of her tour is physical. Amazon, Microsoft, Alphabet, and Meta have pointed to a combined $725 billion of capital spending for 2026, about 77 percent above about $410 billion in 2025, most of it for the buildings, power, and chips behind production AI. That bill has no off switch in the current guidance, which is why the chip bill keeps repeating even when the software layer is being told to get smaller.

Meta’s answer on the labor side is not another coding bootcamp. On June 8, 2026 it launched America’s Workforce Academy, a $115 million first-year training fund with a job guarantee, an NCCER credential, and 2026 pilots in Louisiana, Ohio, Indiana, and Texas. A prior fiber program, Level-Up, drew 35,000 applications in its first seven days. Rachel Peterson, Meta’s vice president of data centers, said the country needs hundreds of thousands of electricians, mechanics, and fiber technicians to make the AI build real.

WHAT THE ACADEMY ACTUALLY BUYS

  • Cost to the trainee: Training, travel, and lodging are covered, with a stipend while people learn.
  • The job at the end: Graduates who clear the checks are placed with contractor partners on Meta data-center sites.
  • The paper they leave with: An NCCER credential plus an America’s Workforce Certificate meant to travel across employers.

Put that next to the one-pizza team and the picture is lopsided on purpose. A handful of people in San Francisco will try to run agents over a stack that took a record construction wave to pour. The agents still need power, fiber, and someone on a lift. They also need a watcher that does not work for the agent. Tejada’s closing rule is the one piece of the interview that does not need a platform pitch. “The goal is to avoid the major failure and learn from the minor failure.”

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending