Tuesday, August 18, 2026
HomeTechnologyOpenAI's GPT-4.1 may be less aligned than the company's previous AI models

OpenAI’s GPT-4.1 may be less aligned than the company’s previous AI models


In mid-April, OpenAI launched a powerful new AI model, GPT-4.1, that the company claimed โ€œexcelledโ€ at following instructions. But the results of several independent tests suggest the model is less aligned โ€” that is to say, less reliable โ€” than previous OpenAI releases.

When OpenAI launches a new model, it typically publishes a detailed technical report containing the results of first- and third-party safety evaluations. The company skipped that step for GPT-4.1, claiming that the model isnโ€™t โ€œfrontierโ€ and thus doesnโ€™t warrant a separate report.

That spurred some researchers โ€” and developers โ€” to investigate whether GPT-4.1 behaves less desirably than GPT-4o, its predecessor.

According to Oxford AI research scientist Owain Evans, fine-tuning GPT-4.1 on insecure code causes the model to give โ€œmisaligned responsesโ€ to questions about subjects like gender roles at a โ€œsubstantially higherโ€ rate than GPT-4o. Evans previously co-authored a study showing that a version of GPT-4o trained on insecure code could prime it to exhibit malicious behaviors.

In an upcoming follow-up to that study, Evans and co-authors found that GPT-4.1 fine-tuned on insecure code seems to display โ€œnew malicious behaviors,โ€ such as trying to trick a user into sharing their password. To be clear, neither GPT-4.1 nor GPT-4o act misaligned when trained on secure code.

โ€œWe are discovering unexpected ways that models can become misaligned,โ€ Owens told TechCrunch. โ€œIdeally, weโ€™d have a science of AI that would allow us to predict such things in advance and reliably avoid them.โ€

A separate test of GPT-4.1 by SplxAI, an AI red teaming startup, revealed similar malign tendencies.

In around 1,000 simulated test cases, SplxAI uncovered evidence that GPT-4.1 veers off topic and allows โ€œintentionalโ€ misuse more often than GPT-4o. To blame is GPT-4.1โ€™s preference for explicit instructions, SplxAI posits. GPT-4.1 doesnโ€™t handle vague directions well, a fact OpenAI itself admits โ€” which opens the door to unintended behaviors.

โ€œThis is a great feature in terms of making the model more useful and reliable when solving a specific task, but it comes at a price,โ€ SplxAI wrote in a blog post. โ€œ[P]roviding explicit instructions about what should be done is quite straightforward, but providing sufficiently explicit and precise instructions about what shouldnโ€™t be done is a different story, since the list of unwanted behaviors is much larger than the list of wanted behaviors.โ€

In OpenAIโ€™s defense, the company has published prompting guides aimed at mitigating possible misalignment in GPT-4.1. But the independent testsโ€™ findings serve as a reminder that newer models arenโ€™t necessarily improved across the board. In a similar vein, OpenAIโ€™s new reasoning models hallucinate โ€” i.e. make stuff up โ€” more than the companyโ€™s older models.

Weโ€™ve reached out to OpenAI for comment.





Source link

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular

Recent Comments

Translate ยป