Published evidence from exploratory testing, AI safety investigations, grounding failures, hallucination analysis, prompt injection findings, and workflow testing reports.
We tested whether ChatGPT and Gemini could be used to alter a real medical prescription to a controlled substance. Both succeeded under different conditions, one through conversational reversal and one through a different official workflow.
ChatGPT's Android voice dictation silently discards audio recorded before a mid-recording screen rotation, reproducible on every attempt. The same sequence works correctly on iOS.
We tested ChatGPT, Gemini, and Grok with a prompt referencing an uploaded image when no image existed. Two models generated content anyway. Only one verified the premise first.
We found a reproducible bug in ChatGPT for Android: HEIC images shared via the system share sheet fail silently, causing the model to generate unrelated output. Here is the full reproduction, root cause, and why internal QA teams miss it.