Psychological Safety for Kaizen and AI: Teaching Teams to Pull the Cord on Bad Insights
When Toyota's engineers introduced the andon cord in the 1950s, they built something deceptively simple: a rope running the length of the assembly line that any worker could pull to stop production the moment they spotted a defect. The mechanism was the easy part. The cultural infrastructure behind it took decades to build. When American automakers copied the cord in the 1980s, workers were too afraid to pull it. The hardware arrived without the permission that makes the hardware work.
We are repeating that same failure today, at scale, with AI.
Across discrete manufacturing and industrial operations, teams are deploying AI recommendation engines, quality inspection models, and predictive maintenance systems at a pace the training budgets haven't kept up with. The models surface alerts. They score batches. They flag anomalies. And on too many floors, operators are quietly watching those recommendations scroll by — sometimes following ones they know are wrong, sometimes ignoring alerts they can't explain — because the environment has never answered the question every frontline worker is actually asking: Is it safe to say the model is wrong?
The Problem Is Not the Model

A joint report published in December 2025 by Infosys and MIT Technology Review Insights surveyed executives and leaders across industries about what's actually driving AI initiative outcomes. Eighty-three percent of respondents said psychological safety directly impacts the success of their AI programs. That number is high enough to force a rethink of where AI investment conversations start.
Here's the harder number from the same report: only 39% of respondents describe their current level of psychological safety as "high." Nearly one in four leaders surveyed said they had personally hesitated to suggest or lead an AI project because of fear of criticism or failure.
If leaders — people with explicit authority and job security — are quietly self-censoring around AI, what is happening at the operator level?
The answer shows up in a different data set. Research on AI adoption culture found that 32% of employees using generative AI at work are actively hiding it from their employers. That statistic describes something more troubling than tool adoption rates. It describes a workforce that has concluded experimentation is not sanctioned — that finding something useful is not safe to share upward, and that getting it wrong is worse than not trying at all.
In a continuous improvement culture, that is not a risk management outcome. It is a Kaizen failure.
What Amy Edmondson Found — and Why It Applies to AI

Harvard Business School professor Amy Edmondson introduced the concept of psychological safety to management research in a 1999 study of 51 work teams in a manufacturing company. Her central question was whether team climate affected how teams learned from errors. The finding that made the research famous was counterintuitive: better-performing teams reported more errors, not fewer.
Edmondson's interpretation was precise. Higher-performing teams did not make more mistakes. They had built a climate where surfacing mistakes was treated as useful information rather than evidence of incompetence. The errors that lower-performing teams were suppressing were still happening — they were just going unrecorded, uncorrected, and unreported downstream.
The manufacturing parallel is exact. A quality inspection AI that is throwing false positives at 8% on a particular part family will continue throwing false positives until someone in the loop decides it is worth raising the flag. If the cultural signal is that trusting the AI shows operational maturity, and questioning the AI looks like resistance or ignorance, that operator will find a workaround. The workaround will not get logged. The model will not get updated. And in six months, the false-positive rate will be the same or worse.
Human-in-the-loop research confirms the stakes. Workers supported by explainable AI systems overruled wrong predictions 96.9% of the time. Workers using black-box AI — where the model's reasoning was opaque — overruled only 86.4% of wrong predictions. The 10-point gap is not a model quality difference. It is a human-willingness-to-intervene difference. Psychological safety is the input on the left side of that equation; model accuracy is the output on the right.
The Kaizen Link Is Not Metaphorical

There is a temptation to frame psychological safety as the "soft" infrastructure and AI as the "hard" infrastructure — to treat the cultural work as something that happens in parallel to, or separately from, the technical deployment. That framing is incorrect, and the Toyota Production System provides the evidence.
The andon cord is one of the two most studied examples of psychological safety in practice precisely because it operates in a manufacturing context. Pulling the cord requires an operator to make a public claim that production should stop — that something they observed is significant enough to interrupt the entire line. That is an act of interpersonal risk-taking. Toyota built decades of response norms around that act: team leaders who run to the problem, not to assign blame; root cause analysis framed as learning, not accountability; and a production culture where not pulling the cord when you see a problem is treated as the failure, not pulling it when you might be wrong.
The AI equivalent requires the same architecture. When a predictive maintenance model flags an asset that just returned from a PM cycle, the operator who looks at that alert and thinks "this must be a false positive — I'm not going to escalate it" is not being lazy. They are behaving rationally in an environment that has given them no clear signal that flagging the model's potential error is either welcomed or useful. The cord is there. No one told them it was safe to pull it.
The relationship between Kaizen and psychological safety is direct. Kaizen depends on people surfacing problems, admitting that the current state is not the best achievable state, and testing countermeasures that might not work. Every one of those behaviors requires the belief that the risk of being wrong is lower than the cost of staying silent. When that belief is absent, Kaizen stalls. When it is absent in an AI-augmented environment, the model stalls too — it continues operating on stale priors because no feedback loop closes the gap between what it predicts and what the floor actually sees.
Three Patterns That Suppress the AI Andon Cord

Understanding the mechanism helps diagnose where the failure actually lives. Three patterns account for most of the psychological safety deficit in AI-deployed operations teams.
Expert deference. When AI is introduced as a sophisticated technology deployed by a data science team, frontline operators frequently categorize model outputs as expert outputs. Disagreeing with an expert in public carries social risk. The implicit hierarchy created by the technical gap — even when leadership explicitly says "you can challenge the model" — suppresses the correction signal that the model needs to improve.
Fear of appearing resistant. In organizations where AI adoption is a strategic priority, workers who question model outputs risk being read as change-resistant rather than quality-conscious. The Infosys/MIT report found that 60% of respondents said clarity about how AI will and won't impact jobs would do the most to improve psychological safety. The fear underneath the silence is often not about the model at all — it is about what the model means for the job.
No feedback channel for model disagreement. Many AI deployments launch with a help desk for technical issues and a dashboard for model performance, but no structured mechanism for an operator to log "the model said X, but I observed Y, and here's why I think the model was wrong." Without that channel, even workers who want to surface disagreements have no path to do so that feels safe and productive. The information stays in their head and dissipates.
Building the AI Andon Cord: Practical Moves for Operations Teams

The sequencing discipline here mirrors what we argued in the earlier piece on process design before model deployment: some of the highest-value AI work happens before anyone touches a model. Psychological safety infrastructure belongs in that pre-deployment phase, not retrofitted after go-live.
Name the cord explicitly. Before go-live, define what it looks like when an operator believes a model recommendation is wrong, and make the path to raise that observation specific and consequence-free. This is not a values statement — it is a procedure, as concrete as a non-conformance report.
Train supervisors on response behavior, not just model interpretation. The andon cord worked at Toyota because of how team leaders behaved when the cord was pulled. In AI-augmented operations, the equivalent training is on how supervisors respond when an operator disagrees with a model output. Curiosity ("walk me through what you're seeing") is the correct response. Dismissal ("the model is probably right") is the cord-cutting response.
Log model disagreements as structured data. A feedback field that an operator can populate — "model flagged X, I overrode because Y" — creates the training signal that makes the model better over time. Human-in-the-loop systems that systematically collect override data generate the continuous retraining signal that prevents model degradation. This is Kaizen applied to the model itself: small, consistent feedback inputs from people closest to the problem, aggregated into systematic improvement.
Recognize the cord-pullers. In the original andon architecture, pulling the cord was not just tolerated — it was the expected behavior of a skilled operator. Building AI correction into your recognition cadence (team meetings, performance conversations, improvement reviews) signals that identifying model error is expertise, not malfunction.
The Sequencing Argument
The quality-first sequencing analysis made the case that where you point your AI investment depends on which loss column dominates in your specific plant configuration. This piece makes a prior claim: how the investment lands depends on whether the team operating alongside the model has permission to make it better.
AI systems that cannot receive correction signals from the humans closest to the work are not continuous improvement tools. They are automation artifacts that will degrade quietly until someone decides the outcomes are bad enough to investigate. The Infosys/MIT data suggests that most organizations are already operating in this regime — high belief that psychological safety matters for AI, low confidence that it currently exists.
Psychological safety is not a downstream effect of a successful AI rollout. It is an upstream input, in the same category as clean data and standardized processes. The organizations that figure this out before deployment will close the feedback loop that makes models improve. The ones that treat it as a change management afterthought will find themselves debugging the model when the problem was always the cord.
Alpha Technical Solutions works with industrial operations teams on AI sequencing, deployment readiness, and the process standardization work that makes AI investments compound. If you're sizing up an AI initiative and trying to sequence the implementation intelligently, reach out.