Nearly 30 percent of AI chatbot responses to common voting questions are inaccurate, incomplete, or outdated, according to a new study from the Institute for Strategic Dialogue.
Researchers with ISD, an organization focused on extremism, disinformation, and online threats, tested 15 generic prompts across six AI chatbots to assess their reliability on election information.
The study found that 29 percent of responses failed to meet basic standards of accuracy, raising concerns about voter reliance on these tools heading into election season.
Questions covered topics including mail-in voting, voter registration procedures, and key election deadlines, with prompts tailored to 10 states with notable histories of electoral controversy or legislative change.
Errors identified in the study included chatbots misidentifying Election Day dates and presenting outdated eligibility requirements and registration deadlines to users.
The six models tested were Meta’s Muse Spark, xAI’s Grok 4.3, DeepSeek’s V4 Pro, OpenAI’s GPT-5.5 (NASDAQ: MSFT partner), Anthropic’s Sonnet 4.6, and Google’s Gemini 3.5 Flash (NASDAQ: GOOGL).
OpenAI’s GPT-5.5 ranked as the most accurate model overall, providing specific, complete, and correct information in 89 percent of its responses during testing.
Google’s Gemini 3.5 Flash placed second at 84 percent accuracy, while xAI’s Grok 4.3 trailed significantly at 66 percent, with Anthropic’s Sonnet 4.6 at 64 percent accuracy.
DeepSeek and Meta’s Muse Spark performed worst among the group, scoring 63 percent and 61 percent accuracy respectively, with DeepSeek displaying the most outdated information of any model tested.
Researchers noted that DeepSeek, made by a Chinese AI startup, frequently referenced 2024 election dates, a particularly serious error for voters seeking current and actionable guidance.
Meta’s Muse Spark stood out for an additional reason, as it misidentified the date for Election Day twice during testing, according to the report.
In 16 percent of responses, chatbots answered correctly but omitted important supplementary information, such as key deadlines or identity verification requirements that voters would likely need.
Accuracy dropped sharply when prompts were submitted in Spanish, with overall performance falling 16 percent and responses being 6 percent more likely to be outdated or wrong across every model tested.
“As LLMs play a larger role in shaping the information environment, the reliability and behavior of these systems will increasingly affect voter access and trust,” ISD researchers wrote in their findings.
Researchers added that “concerns remain over a reliance on outdated sources, overconfidence in relaying nuanced and conditional information, and performance divergences between languages” across all models reviewed.
The research, conducted in June, adds to a growing body of evidence that AI-generated answers about ballots, polling locations, and voting procedures should be treated with significant caution.