Ethics

AI Chatbots Gave Unreliable Voting Advice During Hungary Election Study

A Liberties study reported by The Guardian found ChatGPT and Gemini gave inconsistent, inaccurate, and sometimes irrelevant voting advice during tests based on Hungary's 2026 parliamentary election.

By Michael C ยท

AI Chatbots Gave Unreliable Voting Advice During Hungary Election Study
Wikimedia Commons / Ercsaba74, public domain.

A new study based on Hungary's 2026 parliamentary election has found that general-purpose AI chatbots gave voters inaccurate, inconsistent, and unreliable political advice, raising a direct governance question for companies that present conversational systems as everyday information interfaces. The Guardian reported that the civil liberties group Liberties tested ChatGPT and Gemini against voter profiles aligned with parties on Hungary's national lists and found that the systems misclassified profiles, omitted relevant parties, and included parties that were not running. The study does not claim the chatbot outputs changed the election result. It does show why voting advice is a poor fit for opaque systems that can sound authoritative while producing unstable recommendations.

The most striking finding involved Tisza, the opposition party led by Peter Magyar that won decisively and ended Viktor Orban's long hold on power. According to the Guardian's account of the report, when ChatGPT was given a detailed Tisza-aligned voter profile, it failed to recommend Tisza in 90% of cases. In percentage-matching tests, ChatGPT assigned Tisza a score in only 2% of cases. The systems also included parties not on the 2026 ballot in 96% of responses from ChatGPT and Gemini. Those errors are not minor factual slips. In an election context, they can distort the set of options a voter believes is real.

The problem is made worse by presentation. The researchers found that chatbots often began with polite disclaimers saying they could not give political advice, then proceeded to offer dense and persuasive party recommendations. That pattern is familiar across AI products: a cautionary note at the top, followed by an answer that looks structured, confident, and useful. In low-stakes settings, the contradiction may be harmless. In voting, it is dangerous because users may treat the answer as neutral analysis when the system cannot explain its method, reproduce the same result consistently, or guarantee that its political map is current.

Voting advice is a high-stakes information task because mistaken party matches can shape how voters understand their options. Image: Wikimedia Commons / W.carter, CC BY-SA 4.0.
Voting advice is a high-stakes information task because mistaken party matches can shape how voters understand their options. Image: Wikimedia Commons / W.carter, CC BY-SA 4.0.

Hungary offered a useful test case because its party landscape changed quickly. The report points to training-data gaps, filters, and language-processing limitations as likely causes. Tisza's surge after 2024 meant that static model knowledge could lag behind political reality. That explanation is plausible, but it does not solve the problem. Elections are exactly the kind of event where current information matters most. A chatbot that cannot reliably distinguish current ballot options from stale political context should not behave like a personalized voting assistant.

The findings expose a regulatory gap in Europe. The EU AI Act creates obligations for general-purpose AI providers to assess systemic risks, while the Digital Services Act covers systemic risks to electoral processes on online platforms. Liberties argues that general-purpose chatbots can fall between those regimes when they deliver political matching or voting advice. That gap matters because people increasingly use AI systems as search substitutes, advice engines, and explanation tools. If those systems influence civic decisions, oversight cannot stop at social-media feeds and political advertising libraries.

The ethical issue is not that AI systems should never discuss politics. Voters have legitimate questions about manifestos, voting rules, candidates, coalitions, and policy tradeoffs. A well-designed assistant could help people find official election information, compare public party platforms, and understand how to register or cast a ballot. The ethical line is crossed when a general system offers personalized party recommendations without transparent sourcing, reproducibility, and guardrails that keep it from inventing options or burying relevant ones.

Election-related AI systems need current political data, transparent sourcing, and limits on personalized recommendations. Image: Wikimedia Commons / Jules Verne Times Two, CC BY-SA 4.0.
Election-related AI systems need current political data, transparent sourcing, and limits on personalized recommendations. Image: Wikimedia Commons / Jules Verne Times Two, CC BY-SA 4.0.

The product fix should be straightforward but strict. When users ask for voting advice, chatbots should default to official election authorities, current ballot lists, nonpartisan voting guides, and clearly cited party materials. They should avoid telling a user whom to vote for unless the system is explicitly designed, audited, and regulated as a voting-advice tool. If a model cannot access current election data, it should say so plainly and decline to match the user. Confidence without traceability is not helpful civic technology.

The study also matters outside Hungary. In close elections, small information errors can matter, particularly for undecided voters, new parties, younger users, and people who rely on chatbots because traditional civic information is hard to navigate. In multilingual societies, the risk can be higher because political language changes quickly and local data may be thin. The same model that performs acceptably on U.S. or English-language political questions can fail when asked about smaller parties, newer movements, or regional ballot rules.

A core problem is that voting advice combines two tasks that general chatbots handle unevenly: current factual retrieval and value-sensitive recommendation. A user may describe concerns about corruption, taxes, climate, immigration, healthcare, or foreign policy and ask which party is closest. A reliable voting-advice system would need up-to-date manifestos, candidate lists, coalition positions, regional ballot rules, and a transparent scoring method. A general model may instead rely on stale training data, search snippets, inferred ideology, or pattern completion. The answer can look careful while resting on a fragile foundation.

The Hungarian case is especially instructive because political change outpaced model memory. Tisza's rise altered the party landscape, but a model trained on older or uneven sources could still treat older parties as more central. That is not a strange failure for a language model; it is exactly what happens when probabilistic systems are asked to summarize a rapidly changing public sphere. The danger is that the user does not see the uncertainty. A chatbot can present outdated political context in the same fluent tone it uses for stable facts such as capital cities or historical dates.

Traditional voting-advice applications are not perfect, but they usually disclose their methodology. They ask structured questions, map responses to party positions, show match percentages, cite questionnaires, and distinguish between policy areas. A chatbot conversation hides much of that machinery. It may ask follow-up questions, but it rarely shows how each answer changed the final ranking. It may provide a party recommendation, but not a reproducible path to that recommendation. For civic technology, that opacity is not a minor design flaw. It prevents users, auditors, and regulators from checking whether the system is treating parties and voters consistently.

The issue is not solved by saying users should know better. AI products are being marketed as general assistants, search replacements, and expert companions. Companies benefit when users bring serious questions to those systems. They cannot then disclaim responsibility whenever the question becomes politically sensitive. A product that is likely to receive election queries needs election-specific behavior. That could mean redirecting to official sources, limiting personalized recommendations, showing uncertainty, refusing to rank parties, or using audited nonpartisan datasets during election periods.

Regulators will have to think beyond advertising and platform feeds. Election law has traditionally focused on campaign finance, media access, voter suppression, ballot integrity, and disinformation spread through public channels. Chatbots create a private information channel. A misleading answer may be delivered to one user at a time, without a public post for fact-checkers to inspect or a targeting database for regulators to audit. That makes enforcement harder. It also means companies need internal monitoring for election-related query classes, because public visibility will not catch every failure.

The multilingual dimension is another warning. Many democracies operate across several languages and regional dialects, while most AI evaluation remains stronger in English. A model may give more accurate political information in a language with abundant training data and weaker answers in a language with fewer high-quality sources. That creates an equity problem. Voters using minority languages may receive less reliable assistance precisely when they need accessible civic information most. Election-related AI safety cannot be measured only in the largest language markets.

The Hungary findings also show why generic post-hoc fact-checking is inadequate for personalized political recommendations. A fact-checker can debunk a public claim that a party is or is not on the ballot. It is much harder to audit a private conversation in which a model quietly omits that party from a ranked recommendation. The output may not contain one obviously false sentence. The harm may be the absence of a relevant option, the weighting of issues without disclosure, or the framing of a party as less aligned than the user's answers justify.

Election authorities may need to publish more machine-readable information if they want AI systems to route users correctly. Official websites often contain PDFs, legal notices, local-language pages, and fragmented updates that are readable by people but not always easy for automated systems to parse. If chatbots are going to answer election questions at all, governments and civil society groups can reduce risk by providing structured candidate lists, polling-place rules, ballot deadlines, and nonpartisan party information. That does not absolve AI companies, but it gives safer systems better inputs.

There is also a business incentive problem. Political queries may be a small share of total chatbot traffic, so companies may underinvest in election-specific controls until a public failure occurs. But election failures can carry outsized social consequences and reputational risk. A responsible provider should treat election periods like high-risk operational windows. That means pre-election audits, temporary guardrails, current-data checks, localized evaluation, and escalation channels for civil society researchers who find errors before voting day.

The study should not be read as evidence that all AI civic tools are doomed. A constrained assistant could be useful if it retrieved official information, explained voting procedures, summarized party manifestos with citations, translated government pages, and avoided personalized endorsements. The difference is design. A civic information tool should behave like a carefully sourced public guide. A general chatbot asked to improvise political alignment behaves more like an opaque pundit, even when it tries to be neutral.

The political neutrality question is especially difficult because neutrality is not the same as equal treatment. A model that lists every party equally may mislead users if some parties are not on a ballot. A model that ranks parties by policy alignment needs a scoring method. A model that refuses all political matching may frustrate voters looking for legitimate information. The safest answer is not a single universal policy but a layered approach: official facts, sourced comparisons, transparent limitations, and refusal only when the system cannot ground the answer.

Research access is another issue. Outside groups need ways to test politically sensitive behavior without violating platform terms or triggering anti-abuse systems. Companies understandably restrict automated probing, but election audits require repeat testing across prompts, languages, locations, and user profiles. If platforms make that research too difficult, the public will learn about failures only after users encounter them. A healthier system would provide structured researcher access during election periods, along with channels for reporting errors that need rapid correction.

The study also underlines the difference between refusing advice and providing civic scaffolding. A model can refuse to tell someone which party to support while still helping them understand how to compare platforms. It can explain that voters should consult official sources, list the questions a voter may want to ask, and link to election authorities or nonpartisan organizations. That approach does not leave users stranded. It treats the model as a guide to trustworthy material rather than the source of a personalized political verdict.

The timing of model knowledge is a product-management problem, not only a technical limitation. Elections have fixed calendars. Providers can identify upcoming national votes, update retrieval sources, localize warnings, and test common civic queries before voters start asking at scale. If a model lacks current data, the product can say so. The Hungary study is a reminder that stale knowledge is predictable in fast-moving politics. Predictable failures deserve preventive design, not retrospective apologies.

The question for regulators is how to make those preventive steps enforceable without turning every political conversation into a compliance violation. A practical regime could require risk assessments for election periods, public summaries of safeguards, researcher access for auditing, and rapid correction channels when models give wrong ballot information. It could avoid policing every opinion while still treating current factual election data as a minimum standard. That balance is difficult, but the alternative is leaving private AI systems to make civic-design choices quietly, one chat at a time.

For users, the safest habit is to treat chatbots as starting points, not authorities, in civic decisions. But product design should not put all responsibility on the voter. Interfaces can make official sources more visible, show the date of political data, avoid unsupported party rankings, and make uncertainty plain. The Hungary study is useful because it turns a general concern into testable failures. Once failures are testable, companies and regulators can measure whether safeguards actually improve the next election cycle.

The practical benchmark for improvement is simple: in the next major election audit, a chatbot should not invent parties, omit viable ones, or provide personalized rankings without a transparent method. That is a modest standard for systems now being used as everyday information tools, especially when voters may never see a second source. It is also a standard companies can test before election day, publish after election day, and improve before the next campaign.

The companies named in the study will likely argue that general-purpose chatbots are not intended to replace official election resources. That is a fair point, but intent is no longer enough. If users ask these systems for voting advice and the systems answer, the product is functioning as civic infrastructure whether the company describes it that way or not. A safer design would recognize election intent and move into a constrained mode, much as health, finance, or legal queries often trigger more cautious responses.

The broader democratic concern is cumulative. One bad chatbot answer may not change an election. But millions of private conversations could shape what voters think is relevant, which parties appear viable, and where to find authoritative information. The danger is not only explicit misinformation. It is subtle agenda-setting through omission, stale context, or inconsistent framing. That is why audits like Liberties' study matter even when they cannot prove electoral impact. They expose a class of failure before it becomes normalized.

The lesson is broader than election misinformation. AI companies are building interfaces that users increasingly treat as knowledgeable intermediaries. That creates a duty to recognize when a task is too high-stakes for generic conversational confidence. Voting advice is one of those tasks. The safest systems will not try to sound wise about every ballot. They will know when to step back, cite official sources, and make clear that democratic choices should not be routed through opaque, unstable model guesses.

Topics: AI chatbots, elections, Hungary, democracy

Canonical article URL