Field Guide to AI, Security and Cybercrime

How AI Agents Turn Bad Sources Into Confident Actions

An agent does not just repeat misinformation; it converts it into actions and documents that carry your organisation's authority.

Field Guide to AI, Security and Cybercrime·Abdolmadjid Masoomi·4 October 2026·7 min read

AI agents that search, summarise, and act do not merely present bad information. They operationalise it, turning poisoned web sources into confident decisions, official reports, and automated actions. To manage this, you must enforce source rules and verification steps that simple chatbots never required.

The problem is action, not just information

You are used to the idea that a language model might give you a wrong answer. You ask a chatbot a question, it hallucinates a fact or cites a non-existent source, and you dismiss it. The risk feels contained within the chat window. An AI agent changes that equation completely. An agent is built to take the output of a language model and do something with it: draft an email and send it, summarise a webpage and file a report, analyse a software package and install it. The danger is no longer a wrong piece of text; it is a wrong piece of text that has been stamped with your organisation’s authority and converted into a real-world action. The agent does not just believe the lie; it builds a process upon it.

This shift from passive misinformation to active misoperation is fundamental. A chatbot’s hallucination is a dead end. An agent’s hallucination is the first step in a workflow. It means every unreliable source the agent consults, every fabricated fact it accepts, and every poisoned piece of data it ingests has a direct pathway to becoming a decision, a document, or a transaction. Your defences must now consider not just the truth of a statement, but the consequence of acting upon it.

Why agents trust the wrong sources

Agents typically gather information by searching the web or querying databases. The common, flawed assumption is that the agent applies a human-like sense of source credibility. It does not. An agent uses algorithmic ranking—often the same ranking a search engine provides—which correlates with popularity or search engine optimisation, not with truth or authority. A page that is well-structured, uses the right keywords, and appears high in search results looks credible to the agent’s retrieval system. The language model then summarises that content with confident, authoritative prose, obscuring the rotten foundation.

This creates a perfect environment for poisoning. Malicious actors can craft pages designed specifically to rank for agent queries, packed with plausible-sounding misinformation. These pages are not meant for human eyes; they are traps for the automated retrieval systems that agents rely on. The agent finds the page, treats it as a valid source, and weaves its content into a summary or decision. You are left with a report that cites what looks like a legitimate web address, lending false credibility to a fabricated claim. The problem of fake citations and hallucinations moves from an academic nuisance to an operational hazard.

From wrong belief to irreversible action

The critical failure point is the transition from a wrong belief to a concrete action. An agent’s workflow often follows a simple chain: retrieve information, synthesise it, and execute a command based on that synthesis. There is no inherent circuit-breaker. If the synthesised information states that “vendor X has the best security certification,” the agent can proceed to draft a procurement memo recommending vendor X. If it reads that “package lib-secure-utils version 2.1 is the stable release,” it can trigger an installation script. The action inherits the confidence of the agent’s prose.

Picture an agent tasked with preparing a preliminary supplier assessment. It searches for reviews of a potential IT hardware vendor. It finds a professional-looking review site—a piece of slopsquatting or AI-generated slop designed to rank highly. The site falsely claims the vendor’s factories have superior ethical audits. The agent incorporates this claim into its assessment report, noting the source URL. The report is formatted perfectly, uses corporate branding, and is filed into the supplier evaluation system. A falsehood is now embedded in an official business document, a document that gives the orders for the next stage of due diligence. The scale and speed are transformative; one poisoned source can be converted into hundreds of authoritative reports in minutes.

The specific risk of automated software and commands

The most dangerous manifestation of this is in technical operations. An agent instructed to update software or deploy a tool might search for installation instructions or package names. A poisoned source could recommend a malicious package or a command that compromises system security. The agent, acting on its retrieved instructions, executes the command with the privileges it has been granted. This is not a theoretical risk; the ecosystem of slopsquatting and hallucinated packages exists precisely to exploit automated systems that trust public repositories and forums.

The agent does not know that the command pip install finance-tools-ai is a typosquatting trap. It reads a forum post stating this is the recommended library, retrieves it, and executes the installation. The action is direct and often irreversible the moment it is run. The defence here cannot be based on the agent’s own judgment; it must be based on external rules that prevent such actions outright unless specific, pre-vetted conditions are met.

Controlling the agent: source rules and verification steps

You cannot fix the agent’s ability to judge truth. You must build processes around it that control what it can use and what it can do. This requires moving beyond the capabilities of a chatbot interface to implement agent-specific governance.

First, enforce source allow-lists for consequential tasks. For actions that lead to procurement, legal assessment, or software deployment, the agent should be restricted to a pre-approved list of domains or internal data sources. General web search should be disabled for these workflows. If an allow-list is too restrictive, implement a source tiering system where information from untiered domains triggers a mandatory human verification step before proceeding.

Second, mandate actionable citations. Every factual claim in an agent’s output must be linked to a specific, retrievable source. The link must be to a location a human can actually open and inspect, not an internal vector database ID. This creates an audit trail and makes verification possible.

Third, insert verification before action. The most important control is a hard stop before any irreversible action. The agent should be required to present its planned action and the key sources it relied upon for a human to approve, or for a separate automated system to check against a policy rule. For software installation, this means checking package names against a known-good list or a trusted registry.

Finally, log exhaustively. Every agent decision must be logged with the exact sources that influenced it. This log is not for debugging the model; it is for forensic analysis. When a bad decision is made, you need to answer: which source poisoned the process? This allows you to blacklist domains and understand your attack surface.

Questions people ask

How is this different from a human using a bad source?

A human might also be fooled, but an agent operates at a speed and scale a human cannot match. It can ingest hundreds of sources and produce dozens of reports or actions in the time a human reads one document. It also lacks the subconscious scepticism a person might apply to an unfamiliar website. Most importantly, an agent’s output is often treated as a finished work product, bypassing the human collaboration and discussion that might catch an error.

Can’t we just tell the agent to use reliable sources?

You can instruct it, but you cannot guarantee it. The agent’s concept of “reliable” is statistical, not principled. It will favour what is prevalent and well-linked in its training data or search results. True reliability requires a pre-defined, external rule set—an allow-list—that the agent cannot override with its own reasoning. Instruction is not enforcement.

Do we need to verify every single thing an agent does?

No, you need to risk-stratify its tasks. Low-consequence tasks, like drafting a first summary of public news, might operate with fewer controls. High-consequence tasks, like financial recommendations or system changes, must have stringent source controls and mandatory verification steps. The level of control should be proportional to the potential harm of a wrong action.

Won’t future, more advanced agents solve this problem?

Improved models may hallucinate less and better follow instructions, but the core vulnerability remains: an agent that acts on external data will be vulnerable to poisoning of that data. More capable agents might be entrusted with more consequential tasks, thereby increasing the potential risk. The solution is architectural—building verification and source control into the agent’s workflow—not merely waiting for more reliable models.

Close

The promise of AI agents is automation and scale. Their peril is the automation and scaling of error. When you deploy an agent, you are not deploying a smarter chatbot; you are deploying a system that builds processes on top of whatever information it finds. Your security model must therefore shift from validating answers to validating the sources that feed actions. Implement allow-lists, enforce verification pauses, and maintain detailed logs. The goal is not to make the agent infallible, which is impossible, but to build guardrails that prevent a single bad source from becoming a confident, automated mistake that carries your authority. The agent works for you; you must dictate the rules under which it operates.

Questions people ask

How is this different from a human using a bad source?

A human might also be fooled, but an agent operates at a speed and scale a human cannot match. It can ingest hundreds of sources and produce dozens of reports or actions in the time a human reads one document. It also lacks the subconscious scepticism a person might apply to an unfamiliar website. Most importantly, an agent’s output is often treated as a finished work product, bypassing the human collaboration and discussion that might catch an error.

Can’t we just tell the agent to use reliable sources?

You can instruct it, but you cannot guarantee it. The agent’s concept of “reliable” is statistical, not principled. It will favour what is prevalent and well-linked in its training data or search results. True reliability requires a pre-defined, external rule set—an allow-list—that the agent cannot override with its own reasoning. Instruction is not enforcement.

Do we need to verify every single thing an agent does?

No, you need to risk-stratify its tasks. Low-consequence tasks, like drafting a first summary of public news, might operate with fewer controls. High-consequence tasks, like financial recommendations or system changes, must have stringent source controls and mandatory verification steps. The level of control should be proportional to the potential harm of a wrong action.

Won’t future, more advanced agents solve this problem?

Improved models may hallucinate less and better follow instructions, but the core vulnerability remains: an agent that acts on external data will be vulnerable to poisoning of that data. More capable agents might be entrusted with more consequential tasks, thereby increasing the potential risk. The solution is architectural—building verification and source control into the agent’s workflow—not merely waiting for more reliable models.

Ask NEXUS about this article

Get an AI-powered summary, key points, or follow-up questions about How AI Agents Turn Bad Sources Into Confident Actions, grounded in the essay content and the broader corpus.