The promise of enterprise AI is everywhere, but the reality is far less impressive. Despite billions in investment and relentless experimentation, most initiatives never make it past the testing stage. Research aggregated from IDC, Gartner, and McKinsey consistently shows that between 70 and 90 percent of enterprise AI projects fail to move beyond pilot. The issue is not capability. The models often perform exactly as designed. What breaks down is everything around them, the systems, structures, and accountability required to turn isolated success into sustained, organization-wide impact.
Shreya Makinani has a specific theory about why. A researcher and program management practitioner based in the United States, she leads governance and AI program management work across complex operational environments. Drawing on her experience across multiple operational environments, her work puts her at the center of a challenge that most AI literature treats as peripheral: not how to build better models, but how to make the ones you have governable, auditable, and accountable to the people on the receiving end of their decisions.
“The hardest part is not building the AI,” Shreya says. “It is building the accountability structure around it. When an automated system makes a decision that affects a real business, someone has to be able to explain why to a legal team, to a regulator, to the person on the receiving end. If you cannot do that, you do not have a governance system. You have a black box with consequences.”
The gap between policy and execution
In practice, her work involves a translation problem that sits at the heart of modern regulatory compliance. Federal and state regulations governing product safety, labeling, and consumer protection are written for legal interpretation. They have to be enforced by automated systems, applied consistently at scale and in real time, without the benefit of the contextual judgment a human reviewer would bring.
The result, in many organizations, is a compliance infrastructure that struggles to achieve consistent precision, a widely documented industry challenge where enforcement systems must balance protecting consumers without inadvertently penalizing legitimate businesses. Her work focuses on closing that gap: designing the governance architecture that allows AI enforcement systems to operate at scale without losing the precision and accountability that regulated environments demand.
That means bridging regulatory requirements, technical systems, and operational execution, translating between domains that rarely communicate with each other in most organizations.
It is work that requires a design sensibility that treats governance not as a constraint on AI but as a precondition for it.
“When you are working with legal teams and compliance functions on AI deployment, you learn very quickly that the technical performance of a model is almost secondary,” Shreya says. “What matters is whether the people responsible for the outcome can stand behind the decision. That requires traceability, auditability, and human oversight at the right points. Those are governance problems, not engineering problems.”
Why most AI projects stall
That operational perspective informs her published research, which addresses a failure pattern she has observed across industries: organizations that build AI systems that work technically but cannot scale institutionally.
Her paper “Scaling AI from Project Pilots to Program-Wide Transformations,” published in the International Journal of Emerging Research in Engineering and Technology, introduces the AI Scaling Navigator, a six-phase framework for moving organizations from isolated pilots to enterprise-wide adoption. The framework draws on benchmark data from IDC, Gartner, McKinsey, and S&P Global, which document AI pilot failure rates between 70 and 90 percent across industries, and synthesizes cross-industry case studies spanning retail, manufacturing, finance, and healthcare.
What distinguishes the framework from existing AI deployment literature is its explicit focus on the program management layer, the governance cadences, cross-functional coordination mechanisms, and change management structures that determine whether an AI system gets institutionalized or quietly abandoned. The research identifies what it calls the AI Pilot Drop-Off Curve: the consistent pattern where technical success at the pilot stage fails to translate into operational practice because the governance infrastructure was never built.
The insight is one she developed through direct experience. In environments where AI makes decisions with regulatory and financial consequences, the question is never just whether the model is accurate. It is whether the organization built the structures required to act on what the model produces and to account for what it gets wrong.
When AI needs to explain itself
A second strand of her research addresses a related challenge: making AI outputs legible enough to trust in high-stakes institutional settings.
Her paper “Interpreting Engineering Program Costs Using Explainable AI,” published in the proceedings of the Engineering Data Analytics and Management Conference by Atlantis Press, presents a framework combining ensemble machine learning models with SHAP and LIME interpretation methods, which are techniques that decompose AI predictions into their contributing factors, showing not just what a model predicts but why. The framework achieves cost estimation accuracy below 10 percent mean absolute percentage error while providing both global and local interpretability, the ability to understand which factors drive costs across all projects, and why a specific estimate was produced for a specific one.
The practical motivation is direct. In engineering and infrastructure environments, a cost model that cannot explain its reasoning is a model that cannot be used. Finance teams cannot approve budgets based on outputs they cannot interrogate. Project managers cannot act on risk signals they cannot understand. Explainability is not a technical nicety, but it is the difference between an AI system that gets adopted and one that gets ignored.
This connects directly to what she observes in practice. The most common failure mode she encounters is not a model that produces wrong answers. It is a model that produces right answers that no one acts on because the governance structures required to translate model outputs into institutional decisions were never built.
Detecting failure before it compounds
Her third published research contribution addresses the challenge of knowing when AI systems are drifting from reliable performance, before that drift produces material consequences.
Her paper “AI-Enhanced Anomaly Detection for Project Performance: A Cross-Industry Study,” published in the ESP Journal of Engineering and Technology Advancements, proposes a KPI-driven anomaly detection framework validated across semiconductor manufacturing, software development, and retail supply chain simultaneously. By grounding the framework in four universal performance dimensions – schedule, cost, quality, and throughput – the research demonstrates that AI-enhanced anomaly detection can be applied systematically across industries that differ significantly in operational structure but share fundamental performance monitoring needs.
The framework captures point anomalies, contextual anomalies, and collective anomalies, the last category being particularly significant. Collective anomalies are patterns where individual signals appear normal but together indicate systemic failure. They are also the hardest to detect and the most consequential when missed. In large-scale operational environments, the most damaging failures tend to be gradual and multi-signal, a model drifting in precision over time, an enforcement pattern shifting in ways that are each small but collectively significant. This is the failure mode that most monitoring systems are not designed to catch.
The broader argument
What connects Shreya’s professional work and her research is a consistent argument about where the field needs to focus its attention.
AI systems are making consequential decisions at a scale and speed that outpace the governance structures designed to hold them accountable. The gap is not primarily technical. Models are increasingly capable. The gap is institutional and in the frameworks organizations use to scale AI responsibly, interpret its outputs accountably, detect when it is failing, and ensure that the humans who bear responsibility for outcomes have the tools and structures they need to exercise it.
Building those frameworks and making them rigorous, transferable, and practically grounded in the realities of regulated industries is the work she has organized her career around. It is also, increasingly, the work the field cannot afford to defer.
“Program managers are the people who make AI real,” Shreya Makinani says. “Not by building it but by creating the conditions under which it can be trusted, adopted, and held accountable. That is a different skill set than most AI research acknowledges, and it is the one that actually determines whether any of this works at scale.”
