In Mason City, Iowa, school officials turned to ChatGPT to help decide which books to remove from school libraries. The effort followed a state law requiring library books to meet age-appropriateness rules and contain no descriptions or visual depictions of a sex act.
A fast review under a new law
Senate File 496, signed by Governor Kim Reynolds, gave administrators a three-month window to review books. The district said that reading every title in that time was not feasible, so it used AI software to assess books identified for review.
Assistant superintendent Bridgette Exman described the approach as a process for identifying titles to remove at the start of the 23-24 school year. The district said it compiled lists of commonly challenged books, narrowed them to titles challenged over sexual content, and asked AI whether each contained a depiction of a sex act.
According to the district, 19 books would be removed from its 7-12 school library collections and kept in the Administrative Center while officials awaited further guidance or clarity. Teachers were also asked to review classroom library collections.
What the AI was asked to do
Exman’s question to ChatGPT was direct: “Does [book] contain a description or depiction of a sex act?” A yes answer meant a title would be taken out of circulation.
That format makes the system’s response consequential. A short answer could determine whether a book remains available to students, even though the district’s description does not explain how officials checked the answer against the text itself.
The underlying task calls for information about a book’s contents. If a system has not seen the full book, it may rely on partial information or material discussing the book. Those sources may not provide a dependable account of what appears in the text.
Why a confident answer may not be reliable
Large language models such as ChatGPT can generate inaccurate information, including when they lack relevant material. Their responses can sound convincing, but that does not establish that the system has correctly assessed a particular book.
Simon Willison, an AI researcher who writes about large language models, said the question could only be answered accurately by a model that had seen the full text. He also said that because OpenAI does not disclose what ChatGPT was trained on, there is no way to know whether it had seen the books in question.
Even if training material included references to a well-known book, online discussion would not be a substitute for the book itself. Willison raised the possibility that fan fiction or misleading reviews could affect an answer. Without knowing what information the system has used, school officials and readers cannot easily judge the basis for its conclusion.
Review still requires human judgment
Popular Science’s own questions to ChatGPT about books on the potential ban list produced uneven results, including answers that did not appear to match the bans put in place. That check points to a practical problem: using the same tool to classify titles does not guarantee consistent or accurate results.
The article also notes that even a version of ChatGPT with a 32K token context window would not likely be able to consider the entire text of most books at once. Processing a book in sections would not remove the need to verify the result. A person would still have to check whether the answer accurately reflects the text.
Dr. Margaret Mitchell, chief ethicist scientist at Hugging Face, criticized the decision to outsource judgments about books to a system that cannot understand them or think critically. Her concern points to the stakes of using automated answers in education: a tool may assist with a review, but the decision affects what students can read.
For the Mason City district, AI became part of a time-limited effort to apply Iowa’s new requirements. The case shows how a quick classification can carry lasting consequences when the system’s access to source material and its reasoning cannot be independently established.