BairesDev

AI Agents in Production: Beyond the “Magic Demo”

What changes when AI agents leave the POC: data layers, safe production access, MCP risk, RAG limits, scoped agents. Google Cloud + Carrier panel, 11-min video + transcript.

Last Updated: September 29th 2026
Software Development
23 min read
Jorge Liano
By Jorge Liano
Sr. Google Cloud Practice Director

Jorge Liano is a Senior Google Cloud Practice Director at BairesDev, where he leads cloud and AI initiatives focused on helping organizations design, scale, and operate production-ready solutions on Google Cloud.

illustrative picture of production-ready agentic AI

TL;DR

Most companies can demo an AI agent. Few can run AI agents in production. In a panel with Google Cloud and Carrier, three leaders named what actually changes after the proof of concept: the data layer, the risk tolerance, the access control model, and the size of the agent. The recap video is below, the full transcript is at the end of this page.


Deploying AI agents in production requires four shifts a POC never forces: (1) replace hand-curated demo data with a unified, governed, real-time data layer; (2) treat agent deployment as a risk-tolerance shift and design for confidentiality, integrity, availability and non-repudiation from day one; (3) give agents tool access only through the identity and permissions a human in the same role would inherit; (4) build small, specialized agents with a named owner instead of one “do-everything” agent. When we polled attendees, 44% named trust, reliability and security as their main barrier, twice the next answer. It is not a model problem. It is an architecture and governance problem.

AI agents in production: the 11-minute recap of our panel with Google Cloud and Carrier. The full 52-minute session and the transcript are at the end of this page.

In this recap:

If you have been impressed by AI agent demos, nobody blames you. You watch an agent work through a problem, chain three tool calls together, and produce a result that genuinely beats expectations. Then the demo ends, and the question becomes: how do you operate AI agents in production, reliably, securely, at scale, with real users on the other end?

On December 16, 2025 we hosted “Beyond the ‘Magic Demo’: Getting Agentic AI to Production” with Gio Masawi (Global Solutions Manager, Gemini Enterprise, Google), Jack Lockhart (North America Security Channel Sales, Google Cloud) and Arun Nandi (Chief Data & AI Officer, Carrier). I moderated. The recap above is the 11 minutes that matter most. What follows is the editorial breakdown, then the transcript.

The context is familiar to most teams deploying agents. Gartner expects 33% of enterprise software applications to include agentic AI by 2028, up from under 1% in 2024, and in the same breath predicts that over 40% of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value and inadequate risk controls. Experimentation is everywhere. Production readiness is uneven. The question companies have to answer is whether they have the architectural, operational and governance foundations to run agent systems as part of their core technology stack. As our discussion showed, most teams are not there yet.

What Changes When Agentic AI Leaves the Demo Environment

Proofs of concept are built for speed. You work with clean data, everything is controlled, and there is a high chance it all works. That is fine for testing an idea. Production is a different game.

Arun Nandi sees the same failure repeatedly: POCs rely on manually curated data extracts to demonstrate value quickly, but production agents depend on a unified, governed, real-time data layer to operate across business units. As he put it, “POCs fail to move to scale due to the inability to bring that data fabric and that unified data layer together.”

Gio Masawi’s version is the “happy place” POC. It is small, everything works, and there is an immediate rush to deploy to production. What gets left out is the rigor of testing agent behavior beyond the happy path: refining prompts, adding more data, and gauging whether the agent can handle what production will throw at it.

Jack Lockhart frames it as a risk-tolerance shift rather than a deployment problem. In a POC you are testing a thesis. When the agent hits production, the risks are inherently different, and he often sees teams bolting security on at the end instead of building it in. His reference point is the CIA triad: can you keep confidentiality, guarantee integrity, maintain availability, and track what happened well enough to guarantee non-repudiation?

Once deployed in a production environment, AI agents deal with production data, call critical APIs, execute code, write to downstream systems and interact with real users. They have to meet the same bar as any other enterprise system: reliable, secure, auditable, and cost-predictable. Giving autonomous agents more freedom raises that bar. Which brings us to the foundation of production readiness: trust.

Why AI Agent Deployment Fails: The 44% Problem

We polled attendees live during the session on their main barrier to getting AI agents into production. The results, which are our own first-party data:

  • Trust, reliability and security: 44%
  • Integration with systems and data: 24%
  • Cost and scalability: 16%
  • Organizational readiness: 16%

Jack noted this matched a recent Google Cloud poll in which 76% of respondents put security first. Cost came up strongly there too, because many early deployments were web-facing, where traffic, token consumption and infrastructure costs are far less predictable than for internal agent workloads with known utilization rates.

Read that list against Gartner’s three reasons most agent projects get canceled: unclear business value, inadequate risk controls, and escalating costs once real usage starts. The first item on our poll is the one teams underestimate most, because a demo never tests trust. Production failures rarely come from the model. They come from agent logic that was never exercised outside the happy path, and from enterprise deployments that skipped the operating model.

Trust First: Human in the Loop for Autonomous AI Agents

Agent systems adapt and reason dynamically. You cannot trust them the way you trust a rule-based system. In production, trust comes from having the right controls, the visibility to see what is happening, and knowing who is accountable when something goes sideways.

Keeping a human in the loop is not a temporary fix until the technology improves. It is an intentional design choice for workflows that matter or touch customers directly. It is how you get the benefits of autonomy without giving up responsibility while these systems are still evolving.

What agent failure looks like in the wild

Arun’s example of the largest risk is not a crashed pipeline but a lost customer. He pointed to a car manufacturer whose quoting chatbot was prompted by “an astute individual” into offering a car for one dollar. The damage there is brand equity, and it happens in one interaction. His rule that follows: a human in the loop is critical for every customer-facing use case, with “some wiggle room” for internal operational ones depending on the circumstances.

On the governance side, the teams getting this right treat AI agents like any other risk and compliance issue. Clear boundaries around access, decision authority and escalation create confidence across stakeholders and, counter-intuitively, enable faster progress. But trust and governance only work if agents can reach the right data in the right way.

The Data Layer and Integrations Production Agents Actually Need

AI agents only work if they can talk to your enterprise systems safely and reliably. Integration is where a lot of early excitement hits a wall.

Arun is clear that integration problems are usually data problems in disguise. Your POC ran on clean, curated data that someone hand-picked. In production you need a unified data layer that updates in real time and works across the whole enterprise. Get that right and you cut duplication, governance gets easier, and your agents evolve without rebuilding every connection.

Brittle integrations vs a context convergence layer

The integration mistake Arun sees most is brittleness: point-to-point connections built per agent. His alternative is what he calls a “context convergence layer,” such as MCP (Model Context Protocol) or Google’s A2A. It is an abstraction that carries both the context and the data, so agents become interoperable without custom plumbing, and security and governance are applied to the wrapper rather than to each integration. Scalability comes built in, because a new use case does not mean new plumbing.

Is RAG enough?

We asked the panel how critical it is to connect agents to real-time company data, an inventory database, a product wiki, an internal knowledge retrieval system, and whether retrieval-augmented generation is enough.

Jack’s answer is to start simple. “Simplicity as a V1” is usually the best starting point: achieve the goals you are already confident you can achieve, with a better customer experience, before adding complexity. That lowers risk, shortens development cycles, and keeps token costs from climbing.

Gio’s answer is about evaluation. What teams forget is evaluating the reasoning capacity of the agent: when you get the result set, look at how the agent derived it, which past interactions and data sources it drew on, and where it came from. What he sees instead is continuation, prompting again and again in the hope that an inaccurate response fixes itself without fixing the underlying logic.

Arun’s answer is the line worth remembering: “RAG is not a silver bullet.” It grounds agent outputs in the right data, but if that data is incomplete, inconsistent, or lacks semantic consistency across data sources, hallucinations persist. Better accuracy comes from the right agent architecture and validation mechanisms, not from prompt engineering alone.

Deploying AI Agents Safely: Tool Access, MCP Reach and Layered Defense

The practical question we put to the panel was how to give an agent a tool to reach a production database or a paid third-party API such as Salesforce, and how to stop it from deleting everything.

Inherit the human’s permissions

Gio’s answer is the access control pattern Google uses with Gemini Enterprise: the agent inherits the underlying security of the customer’s platform. If the employee running the agent has access to a given dataset or system, that is the only access the agent inherits. No separate identity model for agents, no reinventing the wheel with the SecOps team.

Reach is the new security risk

Jack’s caution is about reach. MCP servers make it easy for agents to interoperate with external systems, and that is exactly the problem: an agent’s reach “goes a lot further than most APIs would on any one application.” There is no such thing as 100% security, with or without generative AI, so three things need to be in place before any agent deployment:

  1. Attack-surface management. Know what the agent can touch, and keep tool access to the minimum the use case needs. Treat prompt injection attempts as a given, not an edge case.
  2. Detection. Clean logging of tool execution, integrated into the security stack, so unknown issues surface fast.
  3. Response. The ability to recall the power you gave the agent. If you assume a breach will happen, you need the controls to mitigate the damage.

Jack’s summary: if you are going to give an agent power, decide up front how you harness it and how you take it back.

Compliance shapes these choices as much as security does. For high-risk systems the EU AI Act makes automatic event logging a legal requirement, so audit trails of agent decisions and tool execution are not optional in regulated enterprise deployments; they are the evidence you will be asked for.

AI Agent Observability: Running Agent Systems Like Any Production System

Running AI agents in production means seeing what your agents are actually doing, tracing how they made decisions, and stepping in when things go off the rails. If an auditor asks why an agent made a certain decision, you need to show the thought process.

The teams doing this well apply the same operational basics they use everywhere else, logging, monitoring, audit trails and incident response, adapted to how agents work. You need to know what an agent did and why. You need ways to pause, correct or roll back.

Gio’s point is that GenAI ops forces a new set of lenses. Agents are no longer part of the usual tech and ops stack; their capabilities are broad and so is the risk. Agent observability has to run end to end, with the heaviest instrumentation where the risk is highest: agents that interact with customers, and agents with access to sensitive data sets. Most teams have to go back to fundamentals and define a process for how agent ops functions inside the organization, with the right gates before a tested agent gets deployed by three or four enthusiastic engineers.

Jack reinforces this from the security side: detection mechanisms to surface unknown issues, clean logging into security systems, and the ability to intervene. Production requires multiple layers of defense that assume issues will occur.

In practice, AI agent observability covers four things: reasoning traces (what the agent considered), decision audit trails (what it decided and why), tool call logs (which external systems it touched and what it did there), and error rates plus human escalation triggers at critical decision points. OpenTelemetry now publishes semantic conventions for generative AI and agent spans, so trace storage can sit in the same stack as the rest of your production systems instead of a separate agent tool. Continuous monitoring of token consumption belongs on the same dashboard, because cost tracking and cost controls are part of running long-running agents, not an afterthought.

State, memory and what happens when an agent fails mid-task

One difference from stateless services that trips up teams deploying agents for the first time: production agents must maintain context across multiple interactions. Conversation history, intermediate tool results and persistent state all have to survive a restart, a timeout or a model error. If an agent fails mid-task with no checkpoint, it either repeats side effects (a second ticket, a second refund) or silently abandons the work. Design the execution layer so that every tool call is idempotent, every long-running agent checkpoints its state, and a human can resume or cancel from the last known good step.

Small Specialized Agents vs “God Agents”

As AI agents get more powerful, there is a real temptation to build a single agent that can do everything. That is where governance becomes a nightmare.

Gio’s observation is that production-ready teams are not building massive “do-everything” agents. They give each agent a clear responsibility and, where needed, back it with smaller sub-agents. It is less risky to operate, and people adopt it because they know what to expect.

Arun’s warning is about creep. As context windows grow toward infinite, a single agent drifts from read access into write access across more and more systems. Keep agents isolated to their specific use case. He already sees a world of agent proliferation across multiple platforms, where “agent orchestration platforms are a dime a dozen.” A governed ecosystem, a marketplace of sorts where multiple agents are registered and visible, is what keeps shared visibility as they multiply.

This is also the honest case for multi-agent systems. A multi-agent design where specialized agents hand off to other agents through a governed orchestration layer is easier to scale than one autonomous agent with every permission. The agent code stays small, agent behavior stays predictable, and when you need to scale AI agents across business units you add an agent, not a permission.

That is where agent stewardship comes in: a named owner per agent, traceability back to the problem it was built to solve, and a defined escalation path. Without that discipline, agents multiply everywhere and operational complexity spirals. The follow-up question, how to govern agents like employees with their own identity and owner, is the subject of our second panel on agentic AI governance.

Production Readiness Checklist for AI Agent Deployment

Your agent deployment is not production-ready if:

  • Your data lives in siloed, non-API-accessible systems without a unified data layer.
  • Security and access control are planned as a post-launch addition.
  • You have no observability tooling: no logging, no audit trails, no monitoring dashboards.
  • You have not defined who owns each agent and what escalation looks like when it fails.
  • Your agents have no clear, single responsibility; they are designed to do everything.
  • You have not stress-tested beyond the happy path with production-representative data.
  • You have no evaluation harness: no unit tests for tool calls, no scored test set for agent outputs, no regression check in the deployment pipeline.
  • You have no cost controls: no token budget per agent, no alerting on resource consumption, no cost tracking per use case.
  • You have no state model: agents cannot maintain context across interactions or recover when an agent fails mid-task.
  • Your infrastructure provisioning is manual: no managed infrastructure or deployment pipeline that can promote an agent from staging to production with the same controls every time.

Getting one of these right is not enough. Production-ready AI agents require all of them.

Conclusion

Treat AI agents like any other critical enterprise software: proven architecture, operational rigor, and governance adapted for dynamic reasoning. Build in trust, observability and human oversight from day one. The teams moving past the “magic demo” are building agent systems they can operate confidently and trust long term. That is where AI agents deliver real value.

If you are ready to move your agents from experiment to production, our AI software development services team builds and operates production agents with the controls described here, or reach out to our team directly.

Key Takeaways

  1. POCs fail because they rely on clean, curated data. Production agents need unified, real-time data layers across the entire enterprise.
  2. Agent deployment is a risk-tolerance shift. Design for confidentiality, integrity, availability and non-repudiation before launch, not after.
  3. Human in the loop is an intentional design choice for customer-facing and critical workflows, not a temporary workaround.
  4. Give agents only the access the human running them already has. Reach, not leakage, is the new security risk.
  5. RAG is not a silver bullet. Accuracy comes from the data layer and the agent architecture, not from prompting again.
  6. Small, specialized agents beat “god agents”: easier to test, monitor and trust.
  7. Assign agent stewardship: clear owner, clear purpose, clear escalation path. Without it, agents proliferate and complexity spirals.

Watch the Full Session

The full 52-minute session includes the live audience poll, the semantic-layer question, and the closing advice from each panelist.

Transcript: AI Agents in Production Recap (11 min)

Lightly edited for readability. Timestamps match the recap video above. Speakers: Jorge Liano (BairesDev, host), Arun Nandi (Carrier), Gio Masawi (Google), Jack Lockhart (Google Cloud).

The biggest mistake moving an agent from POC to production

[0:00] Jorge Liano: From your perspective, what is the single biggest mistake companies make when they try to move an agent from POC to production? Arun, we’ll start with you.

[0:14] Arun Nandi: For POCs, manually curated extracts are great, because the goal is speed and demonstrating a proof of value. But when you scale into an enterprise, across business units and into production across a much larger fabric, you need a unified, governed, real-time data layer. This is where some POCs fail to move to scale: the inability to bring that data fabric and that unified data layer together.

[0:55] Jorge Liano: Thank you, Arun. Gio, same question.

[0:59] Gio Masawi: From what we’ve seen, most POCs tend to be a happy place. It’s a very small POC, everything works, and there’s an immediate rush to deploy to production. What’s often forgotten and left out is the rigor of testing the agent.

[1:20] Jack Lockhart: I’ll take it from a different point of view. To me it’s less a deployment problem and more a risk-tolerance shift. In a POC you’re testing your thesis. When it hits production, the risks are inherently different. I’d reference the CIA triad: can we keep confidentiality, can we guarantee integrity, can we maintain availability, and are we able to track what is going on to guarantee non-repudiation?

Agent failures in the wild and human in the loop

[1:57] Jorge Liano: What are you seeing as the biggest source of agent failure in the wild, and what’s the right way to implement human in the loop as a safety net?

[2:08] Arun Nandi: From an enterprise perspective, the largest risk is the loss of your company’s brand equity and image with your customers. There’s a great example of an automobile manufacturer that had a quoting agent, a chatbot on one of its websites, which an astute individual prompted into generating a quote of one dollar for a car. Having a human in the loop is going to be critical, especially for customer-facing use cases. There is still some wiggle room for internal operational use cases, depending on the circumstances.

GenAI ops: observability and explaining decisions to auditors

[2:54] Jorge Liano: What does GenAI ops actually look like in practice? If an auditor asks why an agent made a certain decision, how can we possibly show them that thought process?

[3:06] Gio Masawi: This is where observability becomes critical from an end-to-end standpoint, along with building the right processes to manage and monitor our agents, primarily where we’ve got heightened risk: agents that interact with our customers, agents that have access to specific data sets, especially sensitive data. We have to go back almost to fundamental roots and define a process for how we want our agent ops to function within the organization, and make sure the right gates and processes are in place.

Security threats: MCP reach and layered defense

[3:45] Jorge Liano: Jack, from a security standpoint, what is the new threat that keeps you up at night with agents? Is it the agent leaking sensitive data or taking unauthorized actions? What are the guardrails you should be thinking about as you move agents from demos to production?

[4:09] Jack Lockhart: Definitely how well they interoperate with different systems, and how easy that is to accomplish with MCP servers. It’s a little scary, because the reach of that agent goes a lot further than most APIs would on any one application. When advising customers we always say there is no such thing as 100% security, regardless of GenAI or not. So your risk tolerance and your attack-surface management need to be polished. And lastly, response: you need to be able to mitigate the risk. If you’re going to give it power, how do you harness that power and recall it when necessary? There are a lot of layers of defense to put in place so that, if you assume there’s a breach, you have the controls to mitigate the damage.

Giving agents safe access to production systems

[5:12] Jorge Liano: If we get practical: how do we securely give an agent a tool to access a production database, or a third-party API that costs money and holds data, like Salesforce? And how do you stop it from going off and deleting all your data?

[5:35] Gio Masawi: When our customers are leveraging Gemini Enterprise, we inherit the underlying security within the customer’s platform. We work with their security or SecOps team to make sure that if an employee has access to specific data sets or platforms, that is the only access they will inherit once they’re using Gemini Enterprise. So we’re not reinventing the wheel when it comes to security.

Brittle integrations vs a context convergence layer

[6:08] Jorge Liano: Arun, in these integrations, what are the biggest mistakes you see when teams try to connect to systems?

[6:18] Arun Nandi: The challenge is that we build integrations that are too brittle: point-to-point integrations. What you have now is the evolution of what I’d call a context convergence layer, like MCP, which Jack alluded to. Google has put a lot of thought leadership into this with A2A as well. These tackle the challenge head-on by creating an abstraction layer where you have access to the context as well as the data, so you don’t need point-to-point integrations for each agent and the agents become more interoperable. The security and governance aspects are built in, because you’re governing the wrapper. And you get built-in scalability without reinventing the plumbing every time you think of a new use case.

Hallucinations and grounding: is RAG enough?

[7:25] Jorge Liano: How critical is it to connect agents to real-time company data, like an inventory database, a product wiki or an internal system? Is RAG, retrieval-augmented generation, enough, or do we need something more complex?

[7:44] Jack Lockhart: I’ve often seen simplicity as a V1 be the best starting point. Redefining V1 is not about making it more complex, but achieving the goals you’re already confident you can achieve, with a better customer experience. That lowers your risk, delivers immediate results because turnaround times and development cycles are much shorter, and doesn’t push up your costs as significantly.

[8:12] Gio Masawi: What’s often forgotten with our customers is evaluating the reasoning capacity of that particular agent. When you get the result set after prompting or testing your agent, look at how it derived that response. Where is it getting it from? What I often see is just continuation: you prompt, you prompt again, you prompt again, and you’re hoping that based on that initial inaccurate response it fixes itself, without fixing the underlying logic.

[8:52] Arun Nandi: To the points already mentioned, RAG is not a silver bullet. It’s a great start; it grounds responses in the right data. But if your data is incomplete, inconsistent in some manner, or there isn’t semantic consistency, you can still have hallucinations that persist. When you’re looking for better accuracy, that typically comes from the right architecture decisions and mechanisms, not just from prompt engineering. It’s the combination of those factors that leads to better outcomes.

Small scoped agents vs god agents

[9:39] Jorge Liano: Briefly, Arun: for integration, tools and data, what are your thoughts on keeping agents concise and small versus all-knowing?

[9:53] Arun Nandi: As context windows continue to grow toward infinite, there is a risk of creep into more and more areas the agent can touch and influence: not just the original read access, but write access into a number of systems and data. So keep them isolated to their specific use case. I see a world already taking shape with a lot of agent proliferation, spread across multiple systems and platforms; agent orchestration platforms are a dime a dozen today. You’re going to see a lot of proliferation if you open the floodgates. Having a governed ecosystem, a marketplace of sorts where all of these come together, is going to be critical, so that you have shared visibility into how they take shape.

Frequently Asked Questions

  • Three reasons dominate the panel and our audience poll: the POC ran on manually curated data while production needs a unified, governed, real-time data layer; security and governance were planned “later”; and the agent was tested only on the happy path. 44% of attendees named trust, reliability and security as the main barrier, ahead of integration (24%), cost (16%) and organizational readiness (16%).

  • The risk tolerance. A POC tests a thesis under controlled conditions. In production the agent touches live data, critical APIs and real users, so it must meet the same bar as any enterprise system: confidentiality, integrity, availability and non-repudiation, plus predictable cost.

  • For every customer-facing use case, and for any agent that touches sensitive data or takes irreversible actions. Internal operational agents can run with lighter oversight depending on the circumstances. Human in the loop is a permanent design choice, not a temporary fix.

  • No. Retrieval-augmented generation grounds responses in your data, but if that data is incomplete, inconsistent, or lacks semantic consistency across sources, hallucinations persist. Accuracy comes from architecture decisions and evaluating how the agent reasons, not from prompting it again.

  • Let the agent inherit the identity and permissions of the person running it, through the platform’s existing security model, so it can only reach what that person can reach. Add layered defenses, logging of tool execution into your security stack, and a way to recall the agent’s access.

  • Extend standard observability to agent-specific signals: reasoning traces, decision audit trails, tool call logs, error rates and escalation triggers at critical decision points. Instrument most heavily where risk is highest, customer-facing agents and agents with access to sensitive data, and track token consumption alongside.

  • Many small, specialized agents with a named owner. As context windows grow, a single agent creeps from read access into write access across systems. Isolated agents are easier to test, monitor and trust, and a governed marketplace of agents gives shared visibility as they proliferate.

  • Reach. MCP makes it easy for an agent to interoperate with many systems, so a compromised or misbehaving agent can act far beyond what a single API integration could. Treat it with attack-surface management, detection, and a tested response plan.

Jorge Liano
By Jorge Liano
Sr. Google Cloud Practice Director

Jorge Liano is a Senior Google Cloud Practice Director at BairesDev, where he leads cloud and AI initiatives focused on helping organizations design, scale, and operate production-ready solutions on Google Cloud.

  1. Blog
  2. Software Development
  3. AI Agents in Production: Beyond the “Magic Demo”

Hiring engineers?

We provide nearshore tech talent to companies from startups to enterprises like Google and Rolls-Royce.

Alejandro D.
Alejandro D.Sr. Full-stack Dev.
Gustavo A.
Gustavo A.Sr. QA Engineer
Fiorella G.
Fiorella G.Sr. Data Scientist

BairesDev assembled a dream team for us and in just a few months our digital offering was completely transformed.

VP Product Manager
VP Product ManagerRolls-Royce

Hiring engineers?

We provide nearshore tech talent to companies from startups to enterprises like Google and Rolls-Royce.

Alejandro D.
Alejandro D.Sr. Full-stack Dev.
Gustavo A.
Gustavo A.Sr. QA Engineer
Fiorella G.
Fiorella G.Sr. Data Scientist