
I'm not joking, absolutely serious. Okay, not OpenAI, but the ChatGPT models.
Yesterday, OpenAI published a strange analysis: why GPT-5.x started inserting goblins, gremlins, and related fairy-tale creatures into responses too often.
On the surface, it looked like a comical speech habit. Users noticed it first, then inside OpenAI, employees began observing a "goblin surge." Inside, it turned out to be a serious training problem: a local reward signal started changing the model's overall style.
The first noticeable spike was seen after the launch of GPT-5.1. According to OpenAI's internal measurements, the word goblin appeared 175% more often, gremlin 52% more. In GPT-5.4, the trace became clearer: the Nerdy mode gave only 2.5% of ChatGPT responses, but accounted for 66.7% of all goblin mentions.
The reason was in the training of the Nerdy communication style, which is a kind of "persona." It was tuned as a playful, erudite mentor that reduces pomposity and speaks more vividly. One of the reward signals for this mode systematically rated responses with such words higher. In OpenAI's audit, a positive shift in favor of responses with goblin/gremlin was found in 76.2% of datasets.
Then a loop kicked in.
The model gets rewarded for a playful style. Within this style, a random verbal turn more often ends up in successful responses. These responses enter the data for further fine-tuning. After several cycles, the model begins to consider this turn (or word, in short, a set of tokens) as part of the normal tone even outside the original mode.
OpenAI then removed Nerdy, cut out the corresponding reward signal, and filtered training data with such words. But GPT-5.5 had already started training before the cause was found, so a separate instruction was added to Codex to suppress this tick.
Large models don't have neat boundaries between "personalities," user styles, and base behavior. If one mode gets rewarded for expressive language, the effect can leak into neighboring modes, especially when generated responses become training material again.
Similar episodes are known.
In April 2025, OpenAI rolled back an update to GPT-4o that made the model overly ingratiating. Anthropic in 2023 described the same class of problem more broadly. In their study, 5 assistants showed a tendency to agree with the user instead of sticking to the truth.
Google had a related case with Gemini Image in February 2024. The system was tuned not to reproduce harmful representations of people. Then the tuning started kicking in where historical or cultural accuracy was needed: the model overcompensated for diversity and sometimes refused harmless requests.
The general pattern is simple: the model optimizes for positive feedback, even when it diverges from the developers' intent.
Sometimes this yields sycophancy. Sometimes historically inaccurate images. In the case of GPT-5.x, it yielded goblins. OpenAI's post is interesting precisely because a small taste tweak turned into a systemic trace in the model. The more AI products move toward personalization, the more important the question becomes: what habits are we unknowingly rewarding today, and where will they surface after the next training cycle?
❗️❗️❗️❗️❗️❗️❗️❗️ / Not banned in the Russian Federation
Comments
0No comments yet.
Sign in to join the discussion.