Shekar Rao Lakavath is a Principal Service Engineer specializing in enterprise AI, cloud architecture, platform engineering, and AI/ML operations at Providence Health & Services. In this Operator Interview, he makes a consistent case: in a regulated healthcare environment, scaling AI is far less a modeling problem than an operating-model problem — governance, platform foundations, observability, and clear accountability are what move AI from an interesting pilot to a reliable enterprise service.
At Providence you've led enterprise AI, platform, and reliability engineering inside a highly regulated healthcare environment. When you moved your most important AI system from pilot to production, what broke first, and what would you redesign if you started today?
One thing I learned early is that the model usually isn't the first thing that breaks. In a pilot, you're focused on proving value, testing ideas, and getting feedback from a relatively small group of users. Once you move into production, everything around the model becomes the real challenge.
The first friction points were around governance, data access, security controls, privacy requirements, identity management, and operational ownership. Many of those concerns never surface during a pilot because the user base is small and the operational footprint is limited. Once adoption grows, those architectural decisions become highly visible.
In healthcare, you can't simply connect an AI capability to enterprise data and hope for the best. You need clear guardrails around what data can be accessed, who can access it, how outputs are monitored, and how decisions are reviewed when something doesn't behave as expected.
A large part of my work has been helping design the architecture and operational foundations that allow AI to scale responsibly. That includes defining secure integration patterns, supporting privacy-preserving data access, establishing governance controls, building reusable platform services, and making sure AI solutions align with existing security, compliance, and reliability standards. Responsible AI adoption in healthcare requires balancing innovation with security, privacy, and operational trust.
Another lesson was that AI shouldn't be treated as a series of isolated projects. The organizations that scale successfully tend to approach AI as a platform capability. That means investing upfront in common patterns for identity, authorization, observability, model lifecycle management, governance, and risk management. Without those foundations, every new use case creates its own set of operational and compliance challenges.
If I could start over, I would spend even more time on those foundational pieces during the first phase of the journey. I would establish enterprise-wide standards for AI governance, data handling, model evaluation, security, and observability before expanding the number of use cases. I would also bring stakeholders from engineering, security, compliance, legal, and business teams together earlier so ownership and accountability are clear from the beginning.
Looking back, the biggest lesson is that successful AI adoption isn't primarily a technology challenge. It's an operating-model challenge. The organizations that succeed are the ones that build trust, governance, and repeatable platform capabilities alongside the technology itself. That's what ultimately allows AI to move from an interesting pilot to a reliable enterprise service.
Walk us through how you think about the enterprise AI platform in a healthcare setting. What is custom-built versus vendor-provided, and where are the seams that cause the most operational friction?
I tend to think about an enterprise AI platform less as a product and more as an ecosystem. People often focus on the model, but in reality, the model is only one piece of a much larger puzzle.
We generally rely on vendors for things they do exceptionally well, such as cloud infrastructure, foundational AI models, security capabilities, and platform services. Those investments allow organizations to move much faster than if they tried to build everything themselves. At the same time, healthcare organizations have unique workflows, governance requirements, and data considerations that can't simply be solved with an out-of-the-box solution.
That's where custom engineering becomes important. A lot of the work is building connective tissue. How does AI securely access enterprise knowledge? How does it respect data access policies? How do you make sure the right user gets the right information at the right time? How do you evaluate responses, monitor behavior, and apply governance consistently across different use cases?
One thing that surprised me early on is that the hardest part of enterprise AI is usually not the AI itself. Models are becoming increasingly accessible. The real challenge is building an environment where people can actually use AI safely, responsibly, and at scale. In many cases, governance architecture becomes just as important as technical architecture.
Healthcare adds another layer of complexity because most organizations operate across hundreds of interconnected systems, data sources, and business processes that were never originally designed with AI in mind. A significant amount of effort goes into creating trusted pathways between those systems while maintaining privacy, security, compliance, and operational reliability.
The areas that generate the most friction are usually the seams between systems. Identity management, authorization, private networking, data governance, legacy integrations, and observability are often where things become complicated. It's rarely one platform failing. It's usually several platforms interacting in ways nobody anticipated.
I also think we're entering a phase where AI platform engineering is becoming its own discipline. A few years ago, organizations built cloud platforms so development teams didn't have to solve the same infrastructure problems repeatedly. I think we're seeing something similar with AI. The goal isn't to build one successful AI application. The goal is to create reusable capabilities around model access, retrieval, governance, evaluation, observability, and security so teams can innovate without having to reinvent the wheel each time.
At the end of the day, the best enterprise AI platforms don't make AI the center of attention. They make AI feel like a natural extension of how the organization already works, while ensuring the right guardrails are in place for security, privacy, compliance, and trust.
How do you monitor AI systems in production differently from traditional software, and where are your biggest observability gaps in a clinical or regulated context?
Traditional software monitoring focuses on availability, latency, throughput, error rates, dependency health, and resource utilization. Those signals remain essential for AI systems, but they are not sufficient. An AI service can be technically available and return a successful response while still producing an incomplete, inconsistent, poorly grounded, or inappropriate answer.
AI observability therefore needs multiple layers. In addition to infrastructure and application health, teams need to consider retrieval quality, grounding, safety-filter activity, response consistency, model and prompt versions, user feedback, token consumption, cost, and behavior changes over time.
The most significant gaps are often semantic rather than technical. It is relatively straightforward to determine whether an API responded within its latency target. It is much more difficult to determine whether the answer was sufficiently accurate, appropriately contextualized, and safe for the workflow in which it appeared.
A regulated environment introduces another tension. More detailed telemetry can improve troubleshooting and evaluation, but it can also create privacy, retention, and data-minimization concerns. The objective is not to log everything. The objective is to establish meaningful traceability while collecting only what is necessary and appropriately protected.
My research and professional work have reinforced the importance of viewing AI observability as a cross-layer discipline. The evaluation model must connect infrastructure reliability, application performance, AI behavior, governance controls, and user outcomes rather than treating them as separate concerns.
In many ways, observability has evolved from monitoring system behavior to monitoring decision quality and user trust.
Who owns the decision when an AI system produces an outcome that is technically correct but institutionally or clinically risky, and how is accountability structured across engineering, clinical, compliance, and leadership?
One thing I've learned over the years is that AI decisions are rarely owned by a single person or team. Responsible AI deployment is much more of a shared journey.
I like to think about it through the lifecycle of the solution. At the beginning, the conversation is usually about the problem we're trying to solve. Business leaders, operational teams, and subject matter experts help define the use case and what success looks like. As the solution moves into design and development, engineering teams become responsible for building a secure, reliable, and scalable platform. Along the way, privacy, security, compliance, legal, and governance teams help ensure we're designing something that aligns with organizational expectations and regulatory requirements.
As the solution gets closer to production, the discussions become less about technology and more about trust. Depending on the use case, clinical leaders, operational leaders, risk teams, and executive sponsors all have an important role in determining whether the solution is appropriate for the environment in which it will operate.
I've found that one of the most important distinctions is the difference between technical accuracy and organizational readiness. An AI system can generate a response that is technically sound, but organizations still need to ask broader questions. Is the response appropriate for the workflow? Do users understand its limitations? Are there safeguards in place if something unexpected happens?
In my experience, the most successful AI programs create clarity around accountability before a solution reaches production. Everyone understands their role, who owns which decisions, and how concerns are escalated if needed.
At the end of the day, I see enterprise AI as a team effort. The organizations that scale AI successfully aren't necessarily the ones with the most advanced models. They're the ones that bring engineering, security, compliance, operations, and business teams together from the start and keep those conversations going throughout the lifecycle. That's where trust is built, and in healthcare especially, trust is ultimately what enables innovation to scale.
What governance structure do you use to manage AI risk in a regulated healthcare organization, and how has it evolved as your AI footprint has grown?
One thing I've noticed is that governance often gets misunderstood as a control mechanism that's meant to slow innovation. In reality, the most effective governance models are the ones that enable innovation by creating clarity, trust, and repeatable processes. Governance should feel like an accelerator rather than a gatekeeper.
Early in an organization's AI journey, governance is often informal. Teams are experimenting, learning, and trying to understand where AI can deliver value. As AI adoption grows, governance must mature alongside it.
I generally think about governance across four dimensions: use-case governance, data governance, model governance, and operational governance. Together, these create a framework for evaluating risk, managing data responsibly, monitoring model behavior, and ensuring long-term operational ownership.
What has changed most over the last few years is the shift from governing individual models to governing entire AI ecosystems. Organizations are now managing platforms, integrations, retrieval systems, copilots, agents, and increasingly autonomous workflows. Governance has to be embedded into the development lifecycle rather than treated as a review step at the end.
I've also found that successful governance is highly collaborative. Some of the best governance discussions I've participated in were not focused on technology at all. They focused on trust, accountability, transparency, and understanding how people would actually use the solution in practice.
Not every AI application requires the same level of oversight. A productivity assistant helping employees draft content presents a very different risk profile than a system influencing operational decisions or interacting with sensitive information. Organizations that apply a consistent risk-based approach tend to move faster because they can focus the greatest scrutiny where it matters most.
As enterprise AI continues to evolve, I see governance becoming less of a standalone function and more of an operational capability woven throughout the entire AI lifecycle. The goal is not simply to approve AI solutions. The goal is to create a trusted framework that allows organizations to adopt AI responsibly, learn continuously, and scale innovation with confidence.
At what point do you involve legal, compliance, or clinical-safety review in the AI deployment lifecycle, and has that threshold shifted in the last 12 months?
One lesson I've learned is that involving legal, compliance, privacy, security, or clinical-safety teams too late almost always creates unnecessary friction. By the time a solution reaches development or deployment, many of the critical architectural and workflow decisions have already been made. If those stakeholders are only brought in at the end, teams often find themselves revisiting assumptions about data usage, access controls, governance requirements, user interactions, or operational responsibilities.
I prefer to involve those groups early, during use-case definition and initial risk assessment. At that stage, the conversation is less about approving a solution and more about understanding the problem we're trying to solve, the type of data involved, the potential impact of the AI system, and the safeguards that may be needed from the outset. Early collaboration tends to produce better solutions because teams can design with governance in mind rather than adding governance later.
From a lifecycle perspective, I typically see several review checkpoints:
Use-case intake and risk classification — determine the business objective, types of data involved, intended users, level of automation, and overall risk profile.
Architecture and design review — review data flows, security controls, identity management, integration patterns, privacy requirements, and governance considerations before development begins.
Pre-production validation — evaluate model behavior, accuracy, safety controls, prompt management, monitoring capabilities, human-oversight requirements, and operational readiness.
Deployment readiness review — confirm ownership, incident-response processes, auditability, access controls, documentation, compliance requirements, and rollback procedures.
Post-production governance review — monitor adoption, model performance, user feedback, emerging risks, and whether the original assumptions remain valid as the solution evolves.
That said, not every AI application requires the same level of scrutiny. A productivity assistant helping employees summarize documents or draft content has a very different risk profile from a system interacting with sensitive information, influencing operational decisions, or supporting clinical workflows. The review process should always be proportional to the level of risk and potential impact.
Over the last year, I've noticed organizations becoming much more proactive in how they evaluate AI systems. The conversation has evolved beyond simply asking, "Can the model generate a good response?" Today we're asking broader questions: What information can it access? What actions can it take? How is data protected? What happens when the system is wrong? How do we monitor its behavior over time? And perhaps most importantly, do users understand both its capabilities and its limitations?
I think the rise of agentic AI is accelerating that shift. When systems move beyond generating recommendations and begin interacting with enterprise applications or initiating actions, governance can no longer be treated as a checkpoint before go-live. Security controls, approval boundaries, audit trails, human oversight, and rollback mechanisms need to be designed into the solution from the very beginning.
Ultimately, I've found that early engagement creates better outcomes for everyone. It reduces surprises, strengthens trust across teams, and allows organizations to innovate more confidently because the right stakeholders have been part of the conversation throughout the lifecycle rather than being asked to review a finished solution.
What is the single biggest bottleneck preventing your organization from scaling AI beyond its current footprint?
If I had to choose one, I'd say the biggest bottleneck is organizational rather than technical.
Today, most organizations can access powerful models, cloud infrastructure, and AI development tools. Building a proof of concept is often the easy part. The real challenge begins when you're trying to transform a successful experiment into a repeatable, sustainable production capability.
I've seen many organizations successfully demonstrate AI in a controlled environment, but scaling requires much more than technology. It requires trusted data, governance processes, security controls, operational ownership, integration patterns, adoption strategies, funding models, and a clear understanding of who is responsible for long-term outcomes.
One pattern I've observed is that teams tend to approach each AI initiative as a separate project. That works for experimentation, but it doesn't work at scale. Eventually organizations discover they are repeatedly solving the same problems around security, compliance, architecture, monitoring, evaluation, and operational support. That's usually the point where the conversation shifts from building AI solutions to building AI platforms.
For me, scaling AI is really about creating repeatable systems. That means shared intake processes, reusable architecture patterns, common governance controls, evaluation frameworks, observability standards, cost transparency, and well-defined ownership models.
I've also found that success depends on treating AI as a long-term operational capability rather than an innovation project.
Ultimately, I don't think the challenge is launching more pilots. Most organizations already know how to do that. The challenge is creating a reliable operating model that allows innovation to scale without requiring every team to solve the same governance, security, and operational challenges from scratch.
Where has AI genuinely changed how people work day-to-day at Providence? What is the reality rather than the pitch, and where has adoption stalled despite investment?
AI creates practical value when it reduces the time people spend searching, summarizing, organizing, drafting, and navigating fragmented information. In these situations, AI works as an assistive layer that helps employees reach relevant information faster or begin a task more efficiently.
Publicly reported usage of Providence's AI illustrates some of these day-to-day patterns. Early usage included IT and coding questions, writing assistance, proprietary information, and general organizational knowledge. These are practical activities in which AI can reduce friction without independently making a high-impact decision.
Providence's public discussion of AI return on investment also emphasizes measurable operational outcomes rather than only financial projections. Examples described by Providence leadership include supporting employee productivity, reducing provider-message burden, prioritizing inbox content, and improving the specialty-referral process.
However, adoption is not uniform. A technically capable solution can still stall when users do not trust the output, cannot understand its limitations, must perform extensive verification, or do not see how it fits into daily responsibilities. Adoption also slows when the AI experience introduces additional steps or lacks access to the information users need.
The reality is that deployment does not equal adoption. Sustainable use requires role-specific education, transparent communication, visible feedback mechanisms, workflow integration, and continuing measurement. Organizations must be willing to refine or retire use cases based on evidence rather than treating initial launch as the definition of success.
How do you decide what stays human-in-the-loop versus what runs autonomously in a clinical-adjacent system, and how often does that boundary move?
The decision should be based on potential impact, reversibility, uncertainty, data sensitivity, and the consequences of an incorrect action. If an outcome can materially affect a person, a clinical or operational decision, access to information, or another high-impact process, meaningful human review should remain part of the control structure.
Autonomy is more appropriate for tasks that are low risk, bounded, observable, and reversible. Even then, the system should operate within defined permissions, maintain sufficient audit information, and provide an effective way to interrupt or override its actions.
Human-in-the-loop should not mean adding a person as a symbolic final step. The reviewer must receive enough context, time, and authority to make an informed decision. A human reviewer who is overwhelmed by the volume of AI-generated output, or who routinely approves recommendations without meaningful review, is not an effective control.
The boundary can move as evidence accumulates. A use case may begin with close human supervision and later adopt limited automation after its behavior is better understood and reliable monitoring is established. The reverse must also remain possible. If the model, data, workflow, operating environment, or risk changes, the system should return to a more supervised mode.
In healthcare, greater capability does not automatically justify greater autonomy. Evidence, controls, and accountability must progress together.
If you could return to the beginning of your enterprise AI journey in healthcare, what would you do differently in the first 90 days? What is the most common mistake you see peers making when operationalizing AI at scale?
In the first 90 days, I would focus less on maximizing the number of pilots and more on creating the foundations for responsible production use.
I would identify a small set of high-value and measurable use cases, establish a multidisciplinary governance group, define data and risk classifications, select reusable architecture patterns, and create baseline evaluation and observability requirements. I would also clarify who will operate each service after launch and how success will be measured.
Outcome measures should be defined before development begins. Depending on the use case, those measures might include time saved, adoption, response quality, reliability, risk reduction, cost, user confidence, and the amount of human verification still required. Providence's public discussion of AI ROI similarly emphasizes tangible productivity and workflow improvements, including reducing message burden and ineffective referrals, rather than depending only on a single dollar-based calculation.
The most common mistake is assuming that a successful demonstration is almost ready for production. A compelling prototype can conceal unresolved questions about identity, authorization, data quality, integration, evaluation, cost, support, and accountability. The transition from pilot to production is not simply a final deployment step. It is a different engineering and operational discipline.
My broader perspective has also been shaped by publishing research and evaluating technology innovation. In 2026, the International Business Awards named me the Gold Stevie Award winner for AI & Data Leader of the Year; the award program states that its winners were determined from the average scores of more than 300 professionals over two months of judging. Experiences across implementation, research, and industry evaluation have reinforced the same lesson: organizations frequently overestimate the importance of model selection and underestimate the importance of operational discipline.
Note: The views expressed in these responses reflect Shekar Rao Lakavath's personal professional experience and do not necessarily represent the views or policies of his employer.


