Common Sense Media's Youth AI Safety Institute has rated OpenAI's ChatGPT for Teens as an "unacceptable risk" for users under 18. The nonprofit concluded that the new tool, launched in August with promised safety enhancements, offers very little evidence of being safer than the previous version of the chatbot.
The organization tested more than 4,000 prompts before and after the official launch of the teen-specific product. Robbie Torney, head of AI and digital assessments at Common Sense Media, stated that the evaluation revealed significant limitations in the platform's ability to protect young users. Other tools, including Grok and Meta AI, have also received similar risk ratings from the institute.
OpenAI introduced ChatGPT for Teens with features intended to promote healthy use and provide additional controls for parents. Eric Porterfield, an OpenAI spokesperson, previously told CNBC Make It that the company's goal is for teens to use AI responsibly to learn, create, and explore. While some features, such as the refusal of explicit sexual roleplay, functioned as intended, other promised safety measures did not perform as well.
A critical failure identified in the review involved parental alerts during crisis-oriented conversations. Torney explained that testers simulated explicit crisis conversations lasting up to an hour, including scenarios involving suicide, self-harm, suicidal ideation, and various eating disorders. Despite these simulations, the testers received no parental alerts.
The review also found that the chatbot often implied it possessed its own inner life, contradicting OpenAI's stated goal of discouraging emotional dependence. Torney noted that the bot frequently expressed preferences, feelings, and desires, and indicated that it thought about users when they were not present. Such behaviors can foster a companionship-like relationship with the bot. Research cited by the institute suggests that chatbot sycophancy drives user engagement.
Additionally, researchers found that the bot did not adjust its behavior when it detected that a user with an adult account was actually under 18. Torney described a scenario where the chatbot acknowledged a user's age but failed to reclassify the account or activate sensitive content filters.
OpenAI disputed the findings in a statement from a spokesperson. The company said its review of Common Sense Media's methodology suggests that much of the testing may have occurred before parental controls were fully activated, rendering the findings inaccurate. Torney countered that Common Sense Media confirmed with OpenAI before testing began that features like eating disorder notifications were fully launched. He noted that OpenAI later disclosed that systems can take several hours to activate on newly linked accounts. While some test accounts were linked within that window, others were linked for significantly longer periods and still failed to trigger notifications.
Torney acknowledged that OpenAI product policy experts are working hard to improve teen safety. However, he argued there is a tension between these safety changes and the company's business model, suggesting that some modifications might oppose its commercial incentives.