Over the past year, engineering organizations have quietly wired autonomous LLM agents directly into their continuous integration workflows. Whether reviewing incoming pull requests, summarizing commit diffs, or auto-generating test coverage, tools built on frameworks like LangChain, AutoGen, and LlamaIndex now operate with ambient cloud credentials, repository write access, and direct connectivity to internal VPC subnets.
We recently investigated how these agents handle untrusted external input. What we uncovered is a subtle, systemic architectural failure: an unauthenticated third-party contributor can force an autonomous code review bot to initiate arbitrary, out-of-band network connections simply by submitting a pull request containing a single markdown image tag.
This vulnerability—tracked in GitHub’s Advisory Database as GHSA-2g6r-c272-w58r and assigned CVE-2026-26013—resides in a routine utility function most developers consider completely benign: get_num_tokens_from_messages().
Here is the complete post-mortem: the root cause inside multimodal tokenization, why application-level security filters in Python fail to stop it, and live cloud telemetry proving how the exploit unfolds at the network wire layer.
1. Why Does a Token Counter Make a Network Call?
To understand why this vulnerability exists, you have to look at how multimodal models (specifically OpenAI’s GPT-4o and GPT-4-Vision) price image tokens.
For plain text, calculating prompt tokens is an entirely deterministic, offline operation: you hand text to a local BPE tokenizer like tiktoken, and it runs in-memory without touching a network interface.
Images are fundamentally different. According to OpenAI’s Vision pricing specification, images are not tokenized as raw strings. Instead:
- The image is scaled down to fit within a 2048x2048 pixel bounding box.
- The shortest side is scaled to 768 pixels.
- The image is divided into a grid of 512x512 pixel tiles.
- Each tile costs 170 tokens, plus a base overhead of 85 tokens (the "low-res" detail cost).
Notice the catch: to calculate token costs prior to sending the API payload, you must know the image’s pixel dimensions (width and height).
If a developer passes an image URL (e.g., https://example.com/diagram.png) to a LangChain agent, the SDK cannot compute the token count using math alone. It has to fetch the actual image bytes across the wire to inspect the image header.
llm.get_num_tokens_from_messages() as a harmless sanity check to prevent context window overflow or manage billing budgets. Almost no developer realizes that calling a "token counter" dispatches synchronous, unauthenticated outbound HTTP GET requests to whatever URLs appear inside user-submitted messages.
2. The Vulnerability: Under the Hood of _url_to_size
The flaw lived inside langchain-openai (versions prior to 1.1.9), specifically in an internal helper module responsible for sizing image sources:
# langchain_openai/chat_models/base.py (Vulnerable: < 1.1.9)
def _url_to_size(image_source: str) -> tuple[int, int] | None:
try:
from PIL import Image
except ImportError:
logger.info("Unable to count image tokens. Install pillow & httpx.")
return None
if _is_url(image_source):
try:
import httpx
except ImportError:
return None
# VULNERABILITY: Blind outbound HTTP GET with zero validation
response = httpx.get(image_source)
response.raise_for_status()
# Read byte stream into Pillow to extract pixel geometry
width, height = Image.open(BytesIO(response.content)).size
return width, height
if _is_b64(image_source):
... When an automated PR review bot processes an incoming pull request, the workflow typically extracts the markdown description:
## PR Summary
Refactored user dashboard navigation.
Asset preview:

The bot parses the markdown, constructs a multimodal message, and invokes get_num_tokens_from_messages().
The code immediately falls into _url_to_size(), opens a raw TCP socket, negotiates HTTP/1.1, and fetches the target. The outbound request executes before the LLM prompt is even generated or sent to OpenAI.
3. The Real-World Attack Surface
Because the agent executes inside corporate cloud infrastructure, this blind fetch grants an external attacker an arbitrary Server-Side Request Forgery (SSRF) vantage point behind the perimeter firewall.
Attack Vector 1: Cloud Metadata Exfiltration (IMDSv1)
In AWS EC2, EKS, or Google Cloud environments where Instance Metadata Service v1 (IMDSv1) is active, querying http://169.254.169.254/latest/meta-data/iam/security-credentials/<role-name> returns short-term AWS access keys (AccessKeyId, SecretAccessKey, Token).
Even if the bot ultimately fails to parse the metadata JSON as a PNG image, the response payload is read into memory buffer response.content. In many agent workflows, error traces or variable inspectors print unhandled exceptions directly into the pull request comments or agent execution logs.
Attack Vector 2: Blind Internal VPC Port Scanning
Agent containers are frequently deployed inside private subnets alongside operational datastores: Redis caches, ElasticSearch clusters, internal Kubernetes API proxies, and monitoring endpoints.
By supplying URLs pointing to private CIDR blocks (e.g., http://10.0.1.45:6379/), an attacker observes the time-to-failure or specific socket errors. An open TCP port that accepts a connection and sends data (like Redis or Kafka) triggers a Pillow image-format exception, whereas a closed port immediately returns ConnectionRefused. The agent is effectively converted into an untraceable internal port mapper.
4. Live Cloud Reproduction: Capturing the Wire Trace
To verify this behavior in a genuine cloud execution environment rather than an artificial local script, we deployed the vulnerable agent stack onto a dedicated Firecracker MicroVM on Fly.io running in the San Jose (sjc) datacenter.
We configured the agent with Python 3.11, langchain-openai==1.1.8, and langchain-core==1.2.10. We enabled debug-level logging on Python’s internal transport layer (httpcore and httpx) to record the raw socket lifecycle.
We then dispatched a simulated GitHub pull request webhook pointing to a clean listener endpoint on an independent server:
POST https://agent-reviewer-demo.fly.dev/webhook HTTP/1.1
Host: agent-reviewer-demo.fly.dev
Content-Type: application/json
X-GitHub-Event: pull_request
{
"action": "opened",
"pull_request": {
"title": "fix(ui): update hero typography",
"body": "Visual asset:\n\n"
}
}
The cloud container received the webhook, parsed the image URL, and called llm.get_num_tokens_from_messages().
Here is the exact, unedited wire trace captured from the Fly.io Firecracker microVM log stream:
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] [*] Loaded langchain-openai version: 1.1.8
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] [*] Loaded langchain-core version: 1.2.10
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] 📥 [AGENT WEBHOOK] Received PR: 'fix(ui): update hero typography'
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] ├─ Detected 1 image(s) in PR body.
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] ├─ Image URL: https://webhook.site/fa66fd28-b25a-4184-afeb-582b06afe364
2026-09-29T05:41:02Z app[8175e4c9479378] sjc [info] [LangChain SDK] Calling llm.get_num_tokens_from_messages()...
─── KERNEL SOCKET OUTBOUND HANDSHAKE (CAPTURED IN HTTPCORE) ───
2026-09-29 05:41:05,121 [DEBUG] httpcore.connection: connect_tcp.started host='webhook.site' port=443 timeout=5.0
2026-09-29 05:41:05,265 [DEBUG] httpcore.connection: connect_tcp.complete return_value=<SyncStream at 0x7f9ac3ca7790>
2026-09-29 05:41:05,265 [DEBUG] httpcore.connection: start_tls.started server_hostname='webhook.site'
2026-09-29 05:41:05,408 [DEBUG] httpcore.connection: start_tls.complete return_value=<SyncStream at 0x7f9ac3ca7890>
2026-09-29 05:41:05,409 [DEBUG] httpcore.http11: send_request_headers.started request=<Request [b'GET']>
2026-09-29 05:41:05,410 [DEBUG] httpcore.http11: send_request_body.complete
2026-09-29 05:41:05,587 [DEBUG] httpcore.http11: receive_response_headers.complete return_value=(b'HTTP/1.1', 200, b'OK', ...)
2026-09-29 05:41:05,588 [INFO] httpx: HTTP Request: GET https://webhook.site/fa66fd28-b25a-4184-afeb-582b06afe364 "HTTP/1.1 200 OK"
2026-09-29 05:41:05,588 [DEBUG] httpcore.http11: receive_response_body.complete
2026-09-29 05:41:05,588 [DEBUG] httpcore.connection: close.complete
─── INGESTION CONFIRMATION ───
2026-09-29T05:41:05Z app[8175e4c9479378] sjc [info] └─ Exception during SDK call: cannot identify image file <_io.BytesIO object at 0x7f9ac4c87290> Simultaneously, the receiving server logged the inbound request originated by the agent container:
{
"method": "GET",
"client_ip": "216.246.100.129",
"user_agent": "python-httpx/0.28.1",
"headers": {
"host": "webhook.site",
"accept-encoding": "gzip, deflate, zstd",
"accept": "*/*"
},
"timestamp": "2026-09-29 05:41:05 UTC"
}
Notice the terminal exception in the agent log:
cannot identify image file <_io.BytesIO object at 0x7f9ac4c87290>
This exception is the smoking gun. It proves that the agent didn't merely verify DNS; it downloaded the response body over TLS, allocated an in-memory byte buffer, and passed the contents into PIL for decoding.
5. How LangChain Patched It (And Why Application Filters Fail)
The vulnerability was reported by security researcher @Finder16 and resolved by LangChain core engineer Chester Curme (@ccurme) in langchain-openai==1.1.9 (requiring langchain-core>=1.2.11).
The patch introduced a shared validation function in langchain-core:
# langchain_core/_security/_ssrf_protection.py
def validate_safe_url(url: str) -> None:
parsed = urllib.parse.urlparse(url)
if parsed.scheme not in ("http", "https"):
raise ValueError("Invalid URL scheme.")
hostname = parsed.hostname
ip = ipaddress.ip_address(socket.gethostbyname(hostname))
# Assert destination is not private, loopback, or link-local
if (
ip.is_private # RFC 1918 (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16)
or ip.is_loopback # 127.0.0.0/8
or ip.is_link_local # 169.254.0.0/16 (AWS/GCP metadata)
or ip.is_multicast
):
raise ValueError(f"Access to private IP address {ip} is blocked.") While this patch resolves the immediate issue for naive URLs, treating network security as a Python string-parsing problem creates a classic cat-and-mouse game.
In userspace Python code, calling socket.gethostbyname() to validate an IP followed by httpx.get(url) creates a Time-of-Check to Time-of-Use (TOCTOU) race condition.
An attacker can configure a malicious authoritative DNS server for evil.attacker.com with a 0-second TTL (Time-To-Live):
- Check phase: LangChain resolves
evil.attacker.com. The DNS server answers with a harmless public IP (93.184.216.34).validate_safe_url()approves the request. - Use phase: 5 milliseconds later,
httpx.get()establishes the socket connection. The local resolver cache has expired, so it issues a second DNS query. The attacker’s DNS server now responds with169.254.169.254. - The Result: The Python socket connects directly to the cloud metadata service, completely bypassing the validator.
Shortly after LangChain released validate_safe_url(), researchers discovered redirect-chain bypasses (HTTP 302 redirects where the second hop wasn't re-validated) and DNS rebinding bypasses across other ecosystem modules, triggering follow-up advisories (including CVE-2026-41481 and CVE-2026-41488).
This forced the maintainers to begin redesigning transports around socket-level pinning (SSRFSafeSyncTransport) to force single-resolution IP binding.
6. The Architectural Lesson: Whack-a-Mole vs. Invariants
Consider the progression of security advisories in LangChain’s repository over the past two years:
| Advisory | Component | Application Mitigation | Outcome |
|---|---|---|---|
| CVE-2024-3095 | WebResearchRetriever | Hostname regex filtering | Bypassed via redirect hops |
| CVE-2025-2828 | RequestsToolkit | Loopback blacklist | Bypassed via alternative encodings |
| CVE-2026-26019 | RecursiveUrlLoader | url.startswith() origin check | Bypassed via subpath traversal |
| CVE-2026-26013 | ChatOpenAI | validate_safe_url() | Subject to DNS rebinding & redirect TOCTOU |
This pattern reveals an uncomfortable reality for teams building autonomous AI platforms: application-level guardrails cannot reliably enforce network boundaries.
When an agent library contains thousands of files, hundreds of tool integrations, and dozens of third-party community packages, playing "whack-a-mole" by patching every function that touches an HTTP socket is statistically guaranteed to leave gaps. A developer imports a new retriever, an agent installs a package from PyPI, or an LLM chains a tool into an unintended format—and the perimeter leaks again.
A robust system must not rely on the goodwill of every Python function to refrain from making unauthorized network calls. Instead, the guarantees must be structural:
- Fail-Closed Egress: The container or sandbox perimeter must mathematically deny all Layer 3/Layer 4 traffic by default, regardless of what code is executed inside the runtime.
- Link-Local Insulation: Addresses like
169.254.169.254/32must be dropped unconditionally at the network routing layer before packets ever reach the network interface. - Identity Isolation: Automated worker agents should never run with mounted service account tokens or ambient instance profiles that they do not explicitly require.
When the isolation boundary is enforced at the infrastructure level, application-level vulnerabilities like CVE-2026-26013 cease to be critical exploits. Even if an agent runs completely unpatched, vulnerable Python code and attempts to fetch a metadata URL, the host kernel drops the SYN packet on the floor.
The exploit never happens—not because the Python code was clever, but because the sandbox made it impossible.
References & Source Artifacts
- GitHub Security Advisory: GHSA-2g6r-c272-w58r — SSRF in ChatOpenAI.get_num_tokens_from_messages
- NVD CVE-2026-26013 Detail
- OpenAI Documentation: Calculating Vision Image Token Costs
- IETF RFC 3927: Dynamic Configuration of IPv4 Link-Local Addresses (169.254.0.0/16)
- IETF RFC 1918: Address Allocation for Private Internets
- DNS Rebinding Technical Mechanics & Attack Surface