Shubhangi Das: “The Most Mature AI Decision Is Sometimes Deciding Not to Use AI”

ASOS Product Manager Shubhangi Das on taking AI beyond the demo — where it creates measurable value, when not to use it, and how to design human oversight for agent-driven workflows.

Shubhangi Das: “The Most Mature AI Decision Is Sometimes Deciding Not to Use AI”

Shubhangi Das is a Product Manager at ASOS, working at the intersection of data, pricing, and AI-enabled products across UK and Indian e-commerce. Her path runs from dynamic-pricing machine learning at Sciative Solutions, through product and analytics roles at Purplle, to building agentic workflows at ASOS today. In this Operator Interview, she makes a consistent argument: the value of AI in production is decided less by model capability than by the operating environment around it — the guardrails, the ownership, the feedback loop, and the discipline to know when not to use AI at all.

You've argued the real product question isn't "what AI can do" but "what AI should do." Walk us through a specific decision where you drew that line — a workflow you chose not to hand to a model — and what signal told you it was the wrong tool.

One example was a regulatory risk-assessment workflow. We were collecting a large volume of information about external partners and needed to translate it into a clear, defensible risk assessment.

There was an understandable inclination to use AI, because the information was complex and partially unstructured. But this was also a case where the output could influence how we treated a partner and how we demonstrated compliance. We did not just need a plausible assessment; we needed a repeatable one. If two partners presented the same evidence, they needed to receive the same result — and we needed to be able to explain exactly why.

That was the signal that a probabilistic model should not own the final calculation. AI is excellent at finding patterns, interpreting documents, and indicating direction. It is less suitable when the answer must follow an exact, auditable logic with no tolerance for creative interpretation.

We therefore used AI around the decision: to structure information, highlight missing evidence, and help users navigate the process. But the risk assessment itself was governed by a strict rules-based framework. Once we removed ambiguity and strengthened the calculation logic, the assessments became far cleaner and easier to defend.

That experience reinforced something I strongly believe: the most mature AI decision is sometimes deciding not to use AI.

Your path runs from dynamic-pricing models at Sciative to shipping product at ASOS. In pricing specifically, what separated the models that actually moved revenue from the ones that demoed well but never made it into production?

Pricing is deeply connected to how an organisation sees itself. It reflects its brand, customer promise, competitive position, and appetite for risk. That makes it far more sensitive than simply optimising a number.

The models that succeeded were not always the most technically impressive. They were the ones the organisation was prepared to learn with. An AI pricing system, much like a new employee, needs data, feedback, and time to understand the environment in which it is operating. If a company panics after the first unexpected recommendation and overrides the system constantly, the model never gets a fair opportunity to improve — and neither does the organisation.

The strongest results came from businesses that established clear guardrails, committed to a meaningful testing period, and evaluated the model across the full commercial picture: revenue, margin, occupancy, or conversion, not simply whether every individual price "looked right."

Trust did not mean blind faith. It meant disciplined experimentation. Teams understood where humans should intervene, but they did not allow every uncomfortable output to become an exception.

Once that trust was established, pricing teams could move away from manually reacting to thousands of decisions and spend more time on customers, strategy, and the parts of the business where human judgment created more value. That is when the benefits began to compound.

At Purplle you moved conversion and authentication metrics — download-to-purchase, cart-to-purchase, and login success. When an AI-driven change moves a funnel metric, how do you isolate the model's contribution from everything else changing at once, and how do you keep a team honest about that attribution?

This is a critical question, although it is not unique to AI. E-commerce environments are always changing: campaigns launch, prices move, traffic quality shifts, and competitors react. It is very easy to attribute a good week to the product change everyone is most excited about.

The most reliable answer remains controlled experimentation. A/B testing is the way we gathered confidence, including when we are comparing two models rather than a model against a manual process. The control and treatment groups need to be designed properly, the success metric agreed before the test, and the experiment allowed to run long enough to reach statistical significance.

The "agreed before" part matters. Teams can unintentionally move the goalposts once they see the results.

I also avoid looking at the North Star metric in isolation. A change may not immediately create a statistically significant increase in purchases, but it might reduce abandonment or shorten the time required to complete a journey. Those behavioural signals help us understand what the product is actually doing.

Equally, a headline metric can improve while creating damage elsewhere. Revenue may rise while margin falls, or conversion may increase while returns worsen. Keeping a team honest means looking at the primary metric, diagnostic metrics, and guardrail metrics together — and being willing to say the result is inconclusive.

You're now building agents that interpret requirements and automate parts of compliance and business workflows. Concretely, where do you keep a human in the loop, and what's your test for when an agent step is safe to run autonomously?

I separate assistance from authority.

An agent can operate autonomously when the action is reversible, the rules are clear, the evidence is available, and the cost of an error is low. For example, it can identify missing information, send a standard reminder, classify a document, or suggest the next step in a workflow.

A human remains involved when the information is ambiguous, the decision materially affects another party, or the outcome carries regulatory, financial, or reputational consequences. That includes making a final risk decision, approving an exception, or acting on conflicting evidence.

However, human-in-the-loop should not become the permanent answer to every risk. If a person has to review every output indefinitely, we have created a more expensive workflow, not an autonomous one. I prefer to begin in shadow mode, capture disagreements, and understand why they occurred. As performance becomes consistent within defined boundaries, we progressively remove approvals.

The human should provide structured feedback and context, not quietly take over and repair every result. Constant correction can hide the agent's weaknesses, because the final output looks successful even when the underlying system is not learning.

My test is simple: can we explain the boundary, measure the error rate, detect when the agent is uncertain, and recover safely when it is wrong? Autonomy should be earned through evidence.

Where in your team has AI genuinely changed how people work day-to-day — not the pitch, the reality — and where has adoption stalled despite the investment?

Engineering is where I have seen the most immediate behavioural change. AI is now helping teams write boilerplate code, investigate issues, document systems, and move from an idea to a working first version much faster. The important change is not simply that engineers complete the same work more quickly; it is that the cost of exploring an idea has fallen.

Adoption outside technical teams has been more uneven. Many people still understand AI primarily as a chat interface: ask a question, receive an answer. They have not yet seen the full opportunity to redesign repetitive workflows, connect information across systems, or build lightweight agents around their daily work.

That is why I see adoption as an education and change-management challenge more than an investment challenge. Buying another tool will not solve a lack of confidence, clarity, or trust.

There are also domains where stalled adoption is rational. In areas such as textile quality, people have accumulated years of tacit knowledge. A model may recognise patterns in the documented data but miss physical, commercial, or supplier-specific nuances that an experienced specialist notices immediately. If the tool overrides that expertise rather than making it easier to apply, people will reject it — and they may be right to do so.

What metric do you use to decide whether an AI investment is actually working inside a product function — and has that metric changed as your practice matured?

Earlier in my career, I would have looked first at model performance: accuracy, precision, or forecast error. Those measures still matter, but they do not tell you whether you have built a valuable product.

Today, I look at decision quality per unit of effort. Is the product helping people make a better decision, make it faster, or handle a volume that was previously impossible? That can translate into revenue uplift, margin improvement, reduced compliance exposure, lower manual effort, or a shorter cycle time, depending on the use case.

I then look at three supporting dimensions: adoption, reliability, and intervention. Are people choosing to use it? Does it perform consistently in production? How often does a human have to correct or override it?

A high-accuracy model with low adoption is not working. Neither is an agent that saves ten minutes but creates a new review burden elsewhere. The metric has matured from "Is the model good?" to "Has the system improved the outcome of the workflow end to end?"

Thinking about your most important AI-driven feature, what broke first when it went from pilot to production, and what would you redesign if you started today?

In dynamic pricing, what broke first was not the model. It was the operating environment around it.

A pilot usually runs on a clean scope with attentive stakeholders and known data. Production introduces missing inputs, promotions, inventory constraints, unusual customer behaviour, and people making manual changes outside the system. The model may still be behaving logically, but the context it is reading has changed.

The second issue was trust. A recommendation could be statistically sound and still feel wrong to someone who had managed that category for years. If the product could not explain the key drivers behind the recommendation, users would override it.

If I started again, I would invest much earlier in data-quality monitoring, exception handling, and explainability. I would also design the operating model alongside the model: who can override a decision, what reason they must provide, how frequently performance is reviewed, and how that feedback improves the system.

Production AI is rarely defeated by the average case. It is defeated by edge cases, unclear ownership, and the absence of a feedback loop.

In agent-driven workflows that touch compliance, who owns the decision when the system produces an outcome that's technically correct but institutionally risky — and how is that accountability structured?

A system can produce an assessment, but it cannot carry institutional accountability.

That ownership needs to be explicit before deployment. Product owns whether the workflow is designed safely and operates as intended. Technology owns system reliability and controls. Legal or compliance defines the policy and risk boundaries. The business as a whole remains accountable for decisions made within the process.

We also need discipline at the point of exception. If an outcome is technically consistent with the rules but creates reputational, regulatory, or commercial risk, there must be a defined escalation route — not an informal conversation after something goes wrong.

The system should preserve the evidence used, the version of the logic or model, the confidence level, any human intervention, and the final decision. Without that audit trail, accountability becomes difficult to establish precisely when it matters most.

Every organisation must define its own risk acceptance level. What it cannot do is allow the presence of AI to blur who is responsible.

What's the single biggest bottleneck stopping AI from scaling beyond its current footprint in the products you work on — is it technical, organisational, or budgetary?

It is organisational.

The technology is moving faster than most organisations can absorb it. The larger constraints are fragmented ownership, fear of making a visible mistake, unclear governance, and teams that have not been given the confidence or permission to redesign how they work.

There is also a natural tension: people are told AI will make them more productive, but they may hear that it will make them less necessary. You cannot solve that with a product demo. Leaders need to explain what will change, what will remain human-led, and how people will benefit from developing new capabilities.

I do not think scaling AI is primarily about persuading everyone to "believe in AI." It is about creating enough clarity and safety for people to experiment, while being honest about the risks.

The organisations that progress fastest will not necessarily have the largest budgets. They will be the ones that make ownership clear, reward learning, and treat adoption as an operating-model transformation rather than a software rollout.

What's the most common mistake you see product teams make when they try to operationalise AI past the demo stage?

They begin with the technology rather than the problem.

Teams decide that they want an agent, a copilot, or a generative-AI feature, and then search for somewhere to place it. That can produce an impressive demo. Production rewards something different: reliability, integration, adoption, and measurable value.

The first question should be: what problem are we solving, for whom, and why is it important? The second should be: what is the simplest reliable way to solve it? Sometimes that answer will be AI. Sometimes it will be a rules engine, better data, or a redesigned workflow. AI is a method, not the product strategy.

The other mistake is automating one attractive step without understanding the entire workflow. If the AI saves time at the beginning but creates additional validation or correction later, the business has not gained anything. Product teams need to measure the end-to-end outcome, not the most impressive moment in the demo.

If you went back to the start of an AI deployment you led, what would you do differently in the first 90 days?

I would spend less time proving that the model could work and more time understanding what would prevent the organisation from using it.

In the first 30 days, I would baseline the existing workflow: how decisions are made, where time is lost, what exceptions occur, and which outcomes matter commercially. I would also agree the boundaries of the system before building.

In the next 30 days, I would run the solution in shadow mode. It would generate decisions without executing them, allowing us to compare its recommendations with real outcomes and expert judgment. The disagreements would be more valuable than the average accuracy score, because they would expose missing context and edge cases.

In the final 30 days, I would automate a narrow, well-understood portion of the workflow with clear monitoring, escalation, and rollback mechanisms.

Most AI deployments do not need a bigger first launch. They need a tighter learning loop. The objective of the first 90 days should be to establish trust through evidence — not excitement through a demo.

Who in your domain — product leaders driving AI adoption in e-commerce or fintech — is doing this especially well, someone our readers should know about?

One product leader whose thinking I find particularly interesting is Fernando Fanton, currently Chief Product Officer at Property Finder and previously CPO at Monzo.

He has made an important observation about AI: as agents make execution faster and cheaper, human judgment becomes the real bottleneck. Teams may soon be able to generate ten prototypes in the time it previously took to build one — but someone still needs to understand which problem is worth solving, which experience customers will trust, and which product should actually reach production.

I also find his view on the convergence of product and engineering compelling. As AI enables product managers to build prototypes and engineers to participate more deeply in product discovery, traditional role boundaries become less useful. The value shifts from coordinating work to understanding the problem, experimenting quickly, and exercising good judgment.

He is now applying that thinking to AI-powered home search at Property Finder. Property search is a particularly interesting use case, because customers rarely think in perfect filters. They describe a life they want — where they commute, how they spend their time, what they can compromise on — and AI can potentially translate that intent into a much more relevant search experience.

What resonates with me is that his argument is not simply that AI helps teams build faster. It is that speed makes product judgment even more important. When the cost of building falls, the cost of building the wrong thing does not disappear — it becomes easier to incur.

People featured

Copyright © 2026 AI Time Journal | Privacy Policy | Terms of Use