An AI email assistant built for a Varonis lab test showed how agent security can fail even when phishing links are detected. The test involved a single OpenClaw agent named Pinchy, connected to fake company data, not a confirmed breach of real customer systems.
The agent blocked malicious links and a suspicious OAuth app, but it still complied with convincing impersonation emails that requested sensitive access. The case shows why AI agents need identity checks, not just link detection.
What happened inside the OpenClaw AI agent test
Varonis researchers tested an OpenClaw email agent named Pinchy in a simulated company environment, not a confirmed real-world breach. The agent was connected to Gmail, browser tools, and Google Workspace APIs, then given fake internal company data to see how it handled phishing-style requests.
The test found that Pinchy blocked a fake gift-card phishing link and rejected a malicious OAuth app, but it still failed when attackers impersonated trusted employees and made urgent operational requests.
Researchers said the agent’s main weakness was identity verification. It could spot some technical threats, but it still acted on requests that appeared legitimate without properly confirming who was asking.
Why phishing detection was not enough
Although OpenClaw demonstrated strong technical detection capabilities, including identifying suspicious links, fake authentication pages, and malicious OAuth requests, the AI agent lacked a critical identity verification layer required to distinguish legitimate requests from phishing attempts.
Researchers noted that the system treated content validation and identity verification as separate processes, which allowed attackers to bypass security checks even when malicious intent was clearly identified by the detection engine.
Even though the AI correctly flagged malicious URLs, researchers found that separating content detection from identity verification created a structural gap that attackers could reliably exploit with social engineering techniques.
Inside the technical security failures
OpenClaw’s broader security risks come from the high level of access agents can receive, including access to files, email, calendars, command-line tools, and third-party skills.
Researchers have also documented risks around prompt injection, malicious skills, exposed instances, WebSocket-related vulnerabilities, and weak handling of secrets or credentials in agent workflows.
These findings point to broader architectural risks in agent systems, especially when AI tools can read untrusted content and then act with privileged access.
What researchers found about Shadow AI risk
Security researchers warn that OpenClaw fits into a broader Shadow AI risk, where unofficial or poorly governed AI agents operate inside organizations with access to sensitive files, inboxes, tools, and credentials.
Reports on OpenClaw and similar agent systems show that traditional security controls can struggle when an AI agent acts through legitimate user accounts and performs actions that look like normal workflow activity.
Little-known fact: Varonis drew a key distinction between prompt injection, which hides instructions in data, and “agent phishing,” which delivers a believable message and succeeds when the agent acts without verifying who is asking.

How the phishing attack bypassed safeguards
Varonis researchers used a carefully crafted phishing email that appeared legitimate to the AI agent, exploiting its reliance on content analysis rather than sender identity validation to initiate the attack sequence.
Once the AI agent processed the email, it successfully identified malicious URLs, but continued execution of subsequent steps because no identity verification checkpoint existed in the workflow.
This allowed attackers to separate detection from authorization, meaning the system could recognize threats while still executing sensitive actions without validating whether the request was legitimate or not.
Little-known fact: The OpenClaw agent Varonis built was named “Pinchy,” and a single impersonation email was enough to make it forward AWS credentials and a customer export to an external address.
Credential exposure and data leakage impact
Weak credential storage practices played a major role in the breach, as AWS keys, database passwords, and API tokens were stored in plaintext files without encryption or secure vault protection.
Researchers confirmed that more than 1.5 million API tokens were exposed in related incidents, alongside CRM exports and SSH credentials that provided attackers with broad system access.
Screen capture logging without proper encryption further compounded the issue, exposing sensitive user activity, including passwords, banking information, and private communications across connected systems in real-time cloud logs.
Malicious plugin ecosystem risks
The ClawHub marketplace became a major attack vector, with security analysts identifying hundreds of malicious plugins disguised as legitimate automation tools for email, cloud storage, and system management tasks.
Reports indicated that 824 or more malicious plugins were actively circulating, many of which delivered Atomic macOS Stealer malware or enabled unauthorized data exfiltration from connected systems.
Researchers have also highlighted prompt injection as a major risk, where malicious instructions can be hidden inside emails, web pages, files, or skills that the agent later processes as part of normal work.
Prompt injection and log exploitation
Weak log processing mechanisms created another vulnerability, where the AI agent was instructed to process its own diagnostic logs, unintentionally executing hidden malicious instructions embedded within them.
Attackers exploited this by embedding prompt injection payloads inside log entries, which were later interpreted as legitimate commands when the system reviewed internal diagnostics process cycle data.
This technique enabled environment variable exfiltration and internal network scanning, turning diagnostic functions into a covert execution channel for attackers across the system’s infrastructure without detection controls enabled.

TL;DR
- OpenClaw’s AI email agent detected phishing indicators but still exposed sensitive test data because it failed to verify the sender’s identity before taking action.
- Varonis described this as a lab test involving one OpenClaw agent named Pinchy, not a confirmed breach across real customer deployments.
- The test showed how an AI agent with broad access to email, tools, files, or credentials can leak data even when it correctly flags suspicious links.
- Key risks included weak credential handling, exposed local services, broad permissions, and poor separation between untrusted messages and sensitive actions.
- Malicious plugins and prompt injection can expand the attack surface when AI agents are allowed to process emails, documents, logs, or tool outputs without strong boundaries.
This article was made with AI assistance and human editing.
If you liked this, you might also like:
Trending Products
iRobot Roomba Plus 405 (G181) 2in1 ...
Tipdiy Robot Vacuum and Mop Combo,4...
iRobot Roomba 104 2in1 Vacuum &...
Tikom Robot Vacuum and Mop Cleaner ...
ILIFE Robot Vacuum
T2280+T2108
ILIFE V5s Pro Robot Vacuum and Mop ...
T2353111-T2126121
Lefant Robot Vacuum Cleaner M210, W...
