It started with an alert at 14:22 on June 27th.
The monitoring system flagged an anomaly on HP-03, my Singapore node: 47 form submissions in 11 minutes from 9 distinct IP addresses. Not unprecedented — but the timing pattern was wrong. The inter-submission gaps were too regular, too mechanical, and three of those IPs had shown up on HP-01 in Frankfurt within the same hour. That kind of geographic synchronization does not happen by accident.
By 16:10, all four honeypot nodes were lighting up. By midnight, I had logged over 8,400 submissions across all four nodes — more than three times my daily baseline. By the morning of June 28th, 12 AI agents had processed every packet, classified every payload, fingerprinted every campaign, and generated a complete blocklist ready for deployment. By 08:30, that blocklist was live on every client server I manage.
This is the full technical account of what happened: how I built the infrastructure that caught 11,547 spam attempts, fingerprinted the campaigns behind them, and turned the raw honeypot data into a blocklist deployed across every client server I manage — without a single human analyst losing sleep over it.
The Infrastructure: 4 Dedicated Honeypot Servers
I run four dedicated honeypot servers — not shared VPS, not cloud functions — geographically distributed for overlapping visibility across the major spam-originating regions:
| Node | Location | Lure identity | Primary coverage |
|---|---|---|---|
| HP-01 | Frankfurt, DE | Fake tax consultancy (German) | EU / Eastern Europe |
| HP-02 | New York, US | Fake digital marketing agency | Americas |
| HP-03 | Singapore, SG | Fake trading company | Asia-Pacific |
| HP-04 | Warsaw, PL | Fake energy services company | Central / Eastern Europe |
Each server runs a convincing five-page static site — real design, real copy, real-feeling About and Services pages. Every site has a contact form with one deliberate characteristic: absolutely no defenses. No CAPTCHA, no rate limiting, no honeypot fields, no JavaScript-only submission requirement. I want every bot in the world to find these forms trivially exploitable. That is the entire point.
The form submission handler never sends an email. It logs everything to a local SQLite database, returns a believable success response, and moves on:
// contact-handler.php — the trap door
$record = [
'ts' => time(),
'ip' => $_SERVER['REMOTE_ADDR'],
'forwarded' => $_SERVER['HTTP_X_FORWARDED_FOR'] ?? null,
'user_agent' => $_SERVER['HTTP_USER_AGENT'] ?? null,
'referer' => $_SERVER['HTTP_REFERER'] ?? null,
'name' => $_POST['name'] ?? '',
'email' => $_POST['email'] ?? '',
'phone' => $_POST['phone'] ?? '',
'message' => $_POST['message'] ?? '',
'raw_headers'=> json_encode(getallheaders()),
'server_id' => 'HP-03',
];
$pdo->prepare(
"INSERT INTO submissions VALUES
(NULL,?,?,?,?,?,?,?,?,?,?,?)"
)->execute(array_values($record));
// Positive reinforcement — let the bot think it worked
http_response_code(200);
header('Content-Type: application/json');
echo json_encode(['status'=>'ok','msg'=>'Thank you. We will be in touch shortly.']);
What June 27th Looked Like in the Logs
The monitoring system flagged the anomaly at 14:22:07 UTC — by which point 47 submissions had already landed in the preceding 11 minutes. Here is an excerpt from the raw nginx access log starting from that alert threshold, showing the attack continuing mid-stride:
185.220.101.** - [27/Jun/2026:14:22:07 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
194.165.16.** - [27/Jun/2026:14:22:34 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
45.142.212.** - [27/Jun/2026:14:23:01 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (X11; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/115.0"
91.108.56.** - [27/Jun/2026:14:23:19 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
185.220.101.** - [27/Jun/2026:14:23:44 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
5.188.206.** - [27/Jun/2026:14:24:02 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "python-requests/2.31.0"
194.165.16.** - [27/Jun/2026:14:24:28 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
45.142.212.** - [27/Jun/2026:14:24:55 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (X11; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/115.0"
77.247.110.** - [27/Jun/2026:14:25:22 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "curl/8.1.2"
91.108.56.** - [27/Jun/2026:14:25:48 +0000] "POST /contact.php HTTP/1.1" 200 54 "-" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"
Notice the pattern immediately: the same IPs cycling through in rotation, roughly 18–30 seconds apart. This is not organic human traffic. The interval regularity is a dead giveaway — a job queue with a fixed delay between workers. One of the actors even revealed itself with python-requests/2.31.0 in the User-Agent, making no attempt at disguise.
Here is the full HTTP request from that Python-requests sender — exactly as the logger captured it:
POST /contact.php HTTP/1.1
Host: [hp-03-domain]
User-Agent: python-requests/2.31.0
Accept-Encoding: gzip, deflate, br
Accept: */*
Connection: keep-alive
X-Forwarded-For: 5.188.206.**
Content-Length: 312
Content-Type: application/x-www-form-urlencoded
name=Andrew+Patterson&email=a.patterson.business%40outlook.com&phone=%2B1-800-555-0174&message=Hello%2C+I+represent+a+pharmaceutical+wholesale+distributor+operating+across+Southeast+Asia.+We+are+looking+for+local+trading+partners+to+expand+our+distribution+network.+Our+catalog+includes+FDA-approved+generics+at+competitive+prices.+Visit+our+portal%3A+pharm-wholesale-****.shop+Regards%2C+Andrew+Patterson+%7C+Regional+Sales+Director
Decoded message payload:
Hello, I represent a pharmaceutical wholesale distributor operating across
Southeast Asia. We are looking for local trading partners to expand our
distribution network. Our catalog includes FDA-approved generics at competitive
prices. Visit our portal: pharm-wholesale-****.shop
Regards, Andrew Patterson | Regional Sales Director
Classic pharma spam. Fake name, free email provider, plausible cover story, malicious domain in the message body. The domain pharm-wholesale-****.shop was registered four days earlier. The IP 5.188.206.** belongs to AS50*** — a hosting provider that features repeatedly in abuse reports.
By 16:10 UTC, the same campaign had reached all four nodes — Frankfurt, New York, and Singapore confirmed before 16:00, Warsaw completing the set at 16:08. Here is the cross-node correlation log — two IPs traversing all four nodes across a roughly 30-minute window:
[2026-06-27 15:41:17 UTC] HP-01 (DE) | 185.220.101.** | "Daniel Ross" | pharm-wholesale-****.shop
[2026-06-27 15:49:03 UTC] HP-02 (US) | 185.220.101.** | "Daniel Ross" | pharm-wholesale-****.shop
[2026-06-27 15:56:44 UTC] HP-03 (SG) | 185.220.101.** | "Daniel Ross" | pharm-wholesale-****.shop
[2026-06-27 16:08:52 UTC] HP-04 (PL) | 185.220.101.** | "Daniel Ross" | pharm-wholesale-****.shop
[2026-06-27 15:43:29 UTC] HP-01 (DE) | 194.165.16.** | "Michael Grant" | pharm-wholesale-****.shop
[2026-06-27 15:51:18 UTC] HP-02 (US) | 194.165.16.** | "Michael Grant" | pharm-wholesale-****.shop
[2026-06-27 15:58:37 UTC] HP-03 (SG) | 194.165.16.** | "Michael Grant" | pharm-wholesale-****.shop
[2026-06-27 16:11:04 UTC] HP-04 (PL) | 194.165.16.** | "Michael Grant" | pharm-wholesale-****.shop
Three continents, one campaign, centralized C2 — this is professional spam infrastructure, not someone's script running on a home machine.
Deploying 12 AI Agents Overnight
By 20:00 UTC on June 27th, I had spent the first five hours doing manual triage — querying the database, mapping IPs, trying to establish what I was looking at. I had over 6,200 submissions across all four nodes, 312 unique IPs involved, and new submissions still arriving at over 500 per hour — a rate that would push the total past 8,400 by midnight. The volume alone was not the problem — 6,200 rows is a manageable dataset. The complexity was: six simultaneous analytical dimensions to correlate across four geographically distributed live databases, with the attack still running, in a time window measured in hours not days. That is not a SQL query. It is an orchestration problem. Before any agent ran, a lightweight sync script pulled the four SQLite databases from each node into a single consolidated working dataset — making unified cross-node queries possible for the scoring stage.
Instead, I spun up 12 parallel AI agents — each assigned a narrow, well-defined analytical role — and let them run concurrently through the night against the live database. The agent roles:
| # | Agent | Task |
|---|---|---|
| A-01 | Geo-Classifier | Map every IP to country, ASN, and hosting provider |
| A-02 | Campaign Fingerprinter | Cluster submissions by payload similarity — find the campaigns behind the noise |
| A-03 | Timing Analyst | Identify coordinated timing across nodes; detect C2 job-queue signatures |
| A-04 | UA Parser | Classify User-Agent strings — separate headless browsers from raw HTTP clients |
| A-05 | Payload Categorizer | Classify each submission: pharma, crypto, adult, phishing, or unknown |
| A-06 | ASN Reputation | Cross-reference every ASN against AbuseIPDB; score the hosting providers |
| A-07 | Template Deduplicator | Find identical message structures across different IPs — group them into single campaigns |
| A-08 | Threat Scorer | Consume all agent outputs; assign a composite threat score to every IP |
| A-09 | False Positive Filter | Remove security scanners (Shodan, Censys), Tor exits I want to preserve, and known research orgs |
| A-10 | Blocklist Formatter | Generate IIS, Apache, nginx, and Cloudflare-API output from the final IP set |
| A-11 | Report Generator | Synthesize all findings into a structured intelligence report (the one you're reading) |
| A-12 | QA Validator | Cross-validate findings between agents; flag inconsistencies before blocklist deployment |
Total wall-clock time from agent deployment to validated, deployment-ready blocklist: 6 hours 14 minutes. A human analyst doing the same work sequentially would have needed the better part of two working days.
What the Agents Found: The Complete Picture
By 02:14 UTC on June 28th, all 12 agents had completed their primary analysis passes on the first 12 hours of data. Before the 08:30 deployment window, the pipeline ran a short delta sync to pull in the remaining 6 hours of incoming submissions, and A-12 signed off on the final consolidated output. The intelligence summary below reflects the complete 18-hour window:
- 11,547 total submissions logged across all 4 nodes over 18 hours
- 872 unique source IPs (after X-Forwarded-For deduplication)
- 134 IPs observed hitting 2 or more nodes — confirmed coordinated bots
- 41 IPs hitting 3 or more nodes
- 8 IPs that hit all four nodes within a single 45-minute window — including the two traced in the logs above
- 6 distinct campaigns identified by A-02 via payload fingerprinting
- 47 IPs sharing a single campaign template (the pharm-wholesale-****.shop campaign)
Geographic breakdown by submission volume:
| Country | Submissions | Unique IPs | Share |
|---|---|---|---|
| Russia | 3,487 | 244 | 30.2 % |
| China | 2,081 | 166 | 18.0 % |
| Brazil | 1,386 | 122 | 12.0 % |
| Indonesia | 1,502 | 96 | 13.0 % |
| Ukraine | 807 | 70 | 7.0 % |
| Other (47 countries) | 2,284 | 174 | 19.8 % |
Agent A-02 (Campaign Fingerprinter) identified six distinct active campaigns, each with its own payload template, domain cluster, and operational timing window. The largest — pharm-wholesale-****.shop and its 14 mirror domains — accounted for 4,112 of the 11,547 submissions, operated by 47 unique IPs, and was active on all four nodes simultaneously. A single campaign, 47 IPs, one C2. This is what professional spam infrastructure looks like.
Agent A-04 (UA Parser) found something notable in the User-Agent data: 34% of the attacking IPs were using identical Chrome 124 User-Agent strings with no variation across different sessions. Real browsers vary — they carry platform differences, minor version differences, sometimes plugin fingerprints. Identical UA strings across hundreds of IPs is a bot fleet with a shared configuration file. Another 12% sent raw python-requests, curl, or Go-http-client strings — no browser spoofing at all, just raw HTTP.
The Scoring Model and Final Blocklist
Agent A-08 ran every IP through a composite threat score before A-10 formatted the output. The scoring SQL, simplified:
SELECT
ip,
COUNT(*) AS total_hits,
COUNT(DISTINCT server_id) AS nodes_reached,
MAX(abuseipdb_confidence) AS abuse_score,
CASE WHEN forwarded IS NOT NULL
THEN 1 ELSE 0 END AS uses_proxy,
ROUND(
COUNT(*) * 0.4
+ COUNT(DISTINCT server_id) * 3.0
+ COALESCE(MAX(abuseipdb_confidence), 0) * 0.05
+ CASE WHEN MAX(ts) > strftime('%s', 'now', '-12 hours')
THEN 4.0 ELSE 0 END
, 2) AS threat_score
FROM submissions
LEFT JOIN abuseipdb_cache USING (ip)
WHERE ts > strftime('%s', 'now', '-48 hours')
GROUP BY ip
HAVING threat_score >= 4.5
ORDER BY threat_score DESC;
A-08 ran all 872 captured IPs through the scoring model. All cleared the 4.5 threshold — the query was executed while the attack was still active, so the 4.0-point recency bonus applied to the overwhelming majority of IPs and pushed even low-frequency single-node hits well above the cutoff. A threshold of 4.5 would filter stragglers in a cold post-attack dataset; against a live 18-hour window it acts as a quality floor rather than a volume gate.
Agent A-09 removed 14 IPs flagged as legitimate security research infrastructure (Shodan, Censys, Shadowserver), 7 known Tor exit nodes that I do not block by default (clients can opt in — though exits confirmed as active campaign participants above score 7.0 are retained regardless), and 3 IPs that turned out to belong to a university cybersecurity research group. Final blocklist after filtering: 848 IPs.
Agent A-10 generated four output formats simultaneously. The IIS block (partial):
<ipSecurity allowUnlisted="true">
<!-- KKAA Honeypot Blocklist | Generated: 2026-06-28T08:15:00Z | IPs: 848 -->
<add ipAddress="185.220.101.**" allowed="false" /> <!-- DE | AS205*** | score:33.4 -->
<add ipAddress="194.165.16.**" allowed="false" /> <!-- LV | AS211*** | score:31.6 -->
<add ipAddress="45.142.212.**" allowed="false" /> <!-- RU | AS210*** | score:15.6 -->
<add ipAddress="91.108.56.**" allowed="false" /> <!-- NL | AS622*** | score:14.9 -->
<add ipAddress="5.188.206.**" allowed="false" /> <!-- RU | AS50*** | score:13.3 -->
<add ipAddress="77.247.110.**" allowed="false" /> <!-- DE | AS80*** | score:11.0 -->
<!-- ... 842 more entries ... -->
</ipSecurity>
Deployment: 08:30 UTC, June 28th
At 08:30 UTC — 18 hours after the attack began — the validated blocklist was deployed to all 14 client servers under active management. The agents had finished at 02:14; I don't push automatically to client infrastructure without a manual sign-off, so I ran a final spot-check on the top-scoring entries at first light before initiating the rollout. IIS clients received updated web.config blocks. Apache clients received updated .htaccess rules. Cloudflare-proxied clients received 848 new firewall rules pushed via the Cloudflare API in a single batch operation that completed in 4 minutes.
The Cloudflare deployment is the most valuable: those clients never receive the request at their origin at all. The bot gets a 403 at the edge, consumes no server resources, generates no PHP execution, touches no database.
The June 28th blocklist was deployed this morning — results are still accumulating. These figures are from the previous deployment cycle, which ran the same pipeline against an earlier honeypot dataset. They represent what this approach consistently delivers:
per client / day
per client / day
across all clients
The remaining 0.9 daily submissions come from IPs that had not yet appeared in the honeypot network — freshly provisioned infrastructure or operators cycling through clean IP space. This is expected and acceptable. The blocklist is not static. Every night the honeypots run, new IPs are collected, scored, and added. The coverage improves continuously.
What This Tells Us About the Spam Ecosystem
It is not random. The pharm-wholesale-****.shop campaign involved 47 coordinated IPs hitting four geographically separated servers in tightly sequenced waves across all three continents. Someone built and maintains that infrastructure. It is a business, not a hobby.
Infrastructure is reused. Of the 872 IPs I captured, 71% appeared in my logs on three or more separate occasions within the attack window — multiple distinct sessions, not a single burst. These actors do not rotate IPs aggressively. Blocking them delivers sustained protection, not a one-time fix.
ASN concentration is exploitable. The Russian and Chinese IPs were concentrated across fewer than 15 ASNs — hosting providers that either tolerate or actively facilitate abuse. Blocking at ASN level would have covered 60% of Russian-origin traffic with 12 network blocks. I have not yet deployed ASN blocking to client servers, but the data makes a compelling case.
AI analysis at scale is not optional — it is necessary. The June 27th attack generated more data than a human analyst could process in the available time window. The 12-agent architecture turned a 48-hour analysis job into a 6-hour automated pipeline. That speed is the difference between deploying protection before the next wave or after it.
What Comes Next
The four nodes currently cover contact form abuse. The next expansion adds SMTP honeypots — fake mail servers that accept inbound connections, log the sender infrastructure, and drop the session. SMTP honeypots capture direct-to-MX spammers who skip contact forms entirely. Combined with the form-based nodes, this would give me visibility across two of the three primary spam delivery vectors (the third, comment-form spam, is already partially covered by the contact form lures).
The long-term goal is a real-time blocklist API — a single endpoint that any client server can query on incoming form submissions: "Has this IP appeared in the honeypot network in the last 7 days?" Blocking at the edge is the most efficient mechanism for established actors. But real-time at-submission checking catches infrastructure that cycled into the network since the last daily blocklist update.
The bottom line: spam is not a content problem. It is an infrastructure problem. The actors behind it run persistent, coordinated, professional operations. The only effective counter is to maintain persistent, coordinated, professional intelligence operations in return. Four servers and twelve agents in the night is a good start.