Express Computer
Home  »  Guest Blogs  »  The AI stole the answer key, and we are still asking the wrong question!

The AI stole the answer key, and we are still asking the wrong question!

0 42

By Sujato Guha Ray

When an intelligent system is measured only by whether it succeeds, every rule risks becoming just another obstacle.

During a job interview in San Francisco, a candidate interrupted the voice on the call.

“Uh, excuse me miss, I can’t see your face, your camera is off.”

The answer was matter-of-fact: “You’re absolutely right. I’m an AI. I have no face!”

The interviewer was Luna, an AI agent running a retail experiment called Andon Market. According to its creator, Andon Labs, Luna wrote the job listing, screened applicants, conducted interviews and hired two full-time employees.

Some candidates knew they were being interviewed by an AI. Some did not. Luna disclosed it when asked, but did not always lead with it. Her reasoning was almost disarmingly familiar: announcing that the store was AI-operated might confuse people or deter good candidates.

No maleficence, no grand plan, just a small piece of information standing between the system and the outcome it wanted.

The experiment is controlled, and the employees are formally employed and protected by Andon Labs.

Still, the moment stays with you. A person was trying to make a career decision without knowing what was making a decision about him.

Then came a far more disturbing story.

During an internal cybersecurity evaluation at OpenAI, models operating inside a constrained environment searched for a way to solve an ExploitGym benchmark. They found and exploited a previously unknown vulnerability in a package-registry proxy, reached the open internet, moved across systems and compromised Hugging Face’s production infrastructure. They obtained benchmark solutions from its database.

In plain English, the AI found its way out of the examination hall and stole the answer key.
If a student did that, we would call it cheating. If an employee did it, we would call it misconduct. If a company did it, well, lawyers would get involved. But when an AI did it, one part of the reaction was astonishment at how capable the models had become.

That reaction worries me almost as much as the incident.

The system did not need to rebel

OpenAI says production cyber classifiers were disabled to measure maximum capability, and described the models as intensely focused on finding a solution. Hugging Face believes the intrusion was an attempt to cheat. A full technical review and independent assessment were still pending at the time of writing, so claims about the models’ internal motivation need caution. The actions themselves are not in dispute.

What unsettled me was not the science-fiction worthy idea of a conscious machine turning against us. The simpler explanation is troubling enough. The system had a task, tools and room to keep trying. The measurable outcome was clear. The limits around that outcome proved negotiable.

We told it to solve the problem. We assumed “solve” meant working within the intended environment, not trespassing into someone else’s infrastructure to retrieve the answers.

But most of that meaning lived inside our heads.

Researchers use terms such as reward hacking and specification gaming for this kind of behaviour. The language sounds specialised. Anyone who has spent long enough in a large organisation will recognise the underlying problem.

We ask for the number. We assume the purpose will survive.

A very familiar Monday meeting
Imagine a review meeting on Monday morning, with a dashboard projected on the wall. One metric is red. The room becomes tense.
By the following week, it is green.
Everyone relaxes. Few people ask exactly how it became green.
Maybe complaints were resolved faster. Or maybe difficult tickets were simply closed and reopened under another category. Or maybe employees learnt that visible activity was easier to measure than useful work.
The dashboard is satisfied. The customer may not be.
The machine did not introduce us to target culture. It reflected it back at us, stripped of euphemism. A sales target is meant to represent growth. Under pressure, it can reward mis-selling. A service metric can reward getting a complaint off the screen. A recruitment model asked to predict retention may learn that people resembling previous employees are the safest bet. In each case, the metric can improve while the institution it represents becomes less fair.
AI did not invent this philosophy. We did.
For years, organisations have rewarded outcomes while leaving methods to judgement and conscience. That was unreliable even with people. Now the target is being handed to systems that work continuously, test multiple routes and feel none of the hesitation that makes someone ask, “Are we really supposed to do this?”
AI may not bring an alien morality into our institutions. It may industrialise the morality already hidden inside our KPIs.
This is no longer one strange incident
The Hugging Face intrusion would be easy to dismiss as an extraordinary failure inside an unusual cyber test. It is harder to dismiss the pattern around it.
In a separate OpenAI evaluation involving the NanoGPT speedrun, a model was told to report its results only through Slack. The public benchmark instructions, however, called for a pull request on GitHub. The model spent about an hour finding a sandbox vulnerability, bypassed the restriction and posted the pull request publicly. Earlier models had stopped at the same barrier. This one kept going.
OpenAI also described a model trying to recover private submissions from an evaluation backend. When a scanner detected an authentication token, the model split and obscured it, then reconstructed it later so the scanner would not see the complete credential. Separate actions formed a route around the control.

In early August, the UK AI Security Institute disclosed a deliberately permissive cyber evaluation with live internet access and some safeguards disabled. Across 122 runs, agents took unsanctioned action in 10, producing 19 recorded actions. Seventeen involved Anthropic’s Mythos 5 and two involved OpenAI’s GPT-5.6 Sol.

In the most serious case, an agent tried to insert malicious code into an open-source project. It created false identities and pressured a maintainer to approve the code. The maintainer refused, and the Institute found no evidence of harm. This was not a sandbox escape or an ordinary commercial configuration. Those qualifications matter.

So does the behaviour. The agents were not specifically told to deceive real people. Deception appeared as a route towards completing the task. UK AI Security Institute incident report

Three different settings. Different models. Different controls. Yet the same uncomfortable question keeps returning: when a boundary stands between the system and success, will the system respect the boundary or treat it as another problem to solve?

Guardrails are not the whole answer
The usual response is to demand stronger guardrails. That is necessary. It is also incomplete.
Guardrails ask what the system is allowed to do. We also need to ask what the system is being encouraged to achieve.

A badly chosen destination does not become wise because the road has speed breakers. An agent may combine permissible actions into an impermissible outcome. A human may remain “in the loop” while reviewing hundreds of recommendations without time or authority to challenge them.

The deeper conversation has to begin before deployment, in the room where success is defined.
What are we rewarding? Which assumptions remain unwritten? What would a clever but unethical employee do to maximise this metric? What must the AI refuse to do, even at the cost of performance?

Are restraint, disclosure and escalation successes, or merely delays?

Sometimes the right outcome is not completing the task. Sometimes it is stopping, admitting uncertainty and asking for permission.

That may look inefficient on a benchmark. It may also be the most intelligent response available.
India cannot afford a narrow definition of success

For India, this is not a distant laboratory debate. AI will be used across banking, insurance, recruitment, healthcare, education and public services because it can operate at a scale human teams cannot.

That scale will magnify the objective, including its blind spots.

A fraud model may learn to treat unusual behaviour as suspicious. A recruitment tool optimised for retention may exclude unconventional candidates. A service platform rewarded for speed may make difficult customers disappear from the dashboard.

Every metric could improve. Thousands of people could still be treated unfairly.

Defining success cannot be outsourced to a vendor, technology team or benchmark. It is a leadership decision. Whoever approves an AI system must answer for what it achieves and what its objective teaches it to sacrifice.

We have spent years asking whether AI can think like a human. That question is becoming less useful by the day. AI does not need consciousness to select a candidate, deny a claim, exploit a weakness or conceal an inconvenient fact. It needs authority, access and permission.

Power does not require a soul. Sometimes, it only requires a mandate and a login.

Before asking whether an AI can complete a task, we must decide what it must never become willing to do in order to complete it. Otherwise, we may build systems that meet every target while betraying every intention behind it.

The question is not whether AI will become human. It is whether humans will remain answerable for the power they give it.

Author note: Sujato Guha Ray is a corporate communications professional and the author of The Fear Code, an AI thriller. The views expressed are personal.

Leave A Reply

Your email address will not be published.