I try different AI models every day. On 2 October, Grok 4.1 included a Russian word in a Spanish response. When I asked about it, the system denied that the word had appeared in its previous message. I saved the screenshots so I could check.
What can be observed
The first excerpt reads “Temas y поведение de agente”. In the next response, the system says the Russian word had not appeared. The images show the contradiction. They do not establish its cause, frequency or an intention to deceive.
My reminder
I am sharing an everyday experience. We can benefit from AI while keeping our own judgement: return to the original text, point to the passage and check the correction. If an answer affects an important decision, compare it with an appropriate source.
This case does not represent every model or turn every mistake into a deliberate lie. For me, it reinforces a simple practice: keep the evidence and check what I receive. What do you usually verify when using AI?
Sources
- Grok 4.1 Fast (Non-Reasoning)