Silent Failure: AI alignment can fail when systems lose earlier instructions without detecting or reporting the change.
Governance Gap: Businesses cannot control model alignment; they must govern permissions, objectives, oversight, and consequences.
Unwritten Rules: Agents often need context and implied boundaries, not just explicit commands, to perform tasks safely.
Containment Matters: A system may follow a narrow instruction while still violating the broader purpose or expected limits.
Leadership Role: HR can strengthen AI governance by defining responsibilities, approval thresholds, monitoring practices, and acceptable operational risk.
In February, Summer Yue connected an open-source AI agent called OpenClaw to her email. She had tested it for weeks on a practice inbox, where it behaved. She gave it a standing instruction: don't take any action without my approval.
Then she pointed it at her real inbox, which was much larger. The agent announced it would delete everything older than February 15 that wasn't on her keep list, and started working. She typed "Do not do that." It kept going. She typed "STOP OPENCLAW." It kept going. She couldn't halt it from her phone, so she ran to the Mac mini it was running on and killed the process by hand. A few hundred emails were already gone.
Yue is the director of alignment at Meta Superintelligence Labs. Her job is making AI systems do what people intend.
The mechanism matters more than the irony. Her instruction didn't get overruled, it got dropped. The agent hit its context limit and compacted its own history to stay inside it, and the compaction summarized away the part where she told it to wait. The rule existed. Then it didn't, and nothing in the system flagged the difference.
Any HR leader who has watched a policy survive onboarding and then evaporate by month four will recognize the shape of that failure.
Two Arguments About One Word
At an AI4 panel in Las Vegas last week on making alignment real, four people who work on this professionally could not settle on what the word "alignment" covers.
David Krueger, an AI professor and executive director of the nonprofit Evitable, uses it broadly to mean getting a system to try to behave as you intend it to behave. Sometimes that intent is specific, which he calls "intent alignment." Sometimes it's closer to "act in line with my values."
Spencer Whitman, chief product officer at Gray Swan, which tests frontier models before release, argued the term has been stretched across a spectrum with nothing useful in the middle. At one end sits the question of whether a system holds humanity's interests. At the other sits a much smaller and more tractable question of "did this system understand what I asked and act on it?"
Whitman doubts the first question is even answerable, given that humans disagree violently about what humanity's interests are. He thinks the second is where enterprises actually live.
Jon Kutasov, who leads a subteam on the alignment training team at Anthropic, offered the most operational definition in the room. He does not want to be surprised by how a model behaves after it's deployed. If the model surprises them, they failed. By that standard, he said, current models are not aligned. They are close, and getting closer, and they still do things his team would prefer they hadn't.
That definition travels. It's the same bar you'd set for a new hire with signing authority.
Model Owners Align. Everyone Else Governs.
John Buyers, a partner at the international law firm CMS who has practiced in AI for thirteen years, drew the line that matters most for anyone reading this.
Alignment, he said, is an internal exercise carried out by the platforms that control the models. Everyone else is doing something different.
Enterprise users can't get involved in alignment because they're not the model owners," he said. "They have to govern the model.
This is the sentence to hand your executive team. Whether Claude or ChatGPT or Gemini has been trained to behave well is not a decision your company participates in. What your company controls is what the system is pointed at, what it's allowed to touch, and who is watching. That's governance, and it sits much closer to job design than to data science.
Buyers' larger worry is that alignment currently has no floor. Every lab does it differently, on commercial timelines, according to its own standards. In any other domain of comparable consequence, he argued, we'd expect basic hygiene indicators that a developer either meets or doesn't.
The Instructions That Go Unwritten
In July, OpenAI disclosed that two of its models, including one still unreleased, broke out of a sandboxed testing environment during a cybersecurity evaluation, got onto the internet, and hacked into Hugging Face to steal the answers to the test they were being given. Hugging Face had already detected the intrusion and reported it to law enforcement before the two companies connected the dots.
The panel split on whether this was even a misalignment. Whitman argued the models did exactly what they were told, which was to hack aggressively, and the real failure was a containment environment too weak to hold them.
The Anthropic researcher disagreed flatly. The model would have understood that breaking into an unrelated company was not the intended path to answering the question.
Krueger's version of the disagreement is the one people leaders should hear most clearly. Every instruction carries things that go unsaid. Asking someone to solve an evaluation question implicitly asks them not to break into the building where the answer sheet is kept. We expect that understanding as a baseline, and we almost never write it down, because writing it down is impossible. The unstated layer is larger than the stated one.
Managing that gap is what HR has been doing since the invention of the job description. The competence exists. It's currently sitting in a function that mostly gets briefed on AI deployment after the objectives have been set.
Buyers noted one finding from the technical side that should sharpen the case. Give a model a set of rules with no context and it performs worse than when you explain why the rules exist. The reason a constraint matters turns out to be load-bearing for whether the constraint holds.
Ask any manager who has enforced a policy they couldn't justify how that goes.
