Teaching the Robots to Read My 37 Domains
I own 37 domains. This is not a brag. This is a confession. Nobody needs 37 domains. I have them the way some guys have project cars up on blocks — except my project cars all have DNS records and cron jobs, and they judge me.
This week I decided the robots should be able to read them. All of them. Search engines, crawlers, the whole indexing apparatus. Which meant sitemaps. Which meant I had to answer a question I'd been dodging for years: what pages do I actually have?
I did not know. Read that again. I run this estate. I have an AI chain of command with a Primary and subordinates and standing orders, and if you'd asked me to list every URL across my own domains I would have gotten maybe 60% and then started guessing like a lance corporal at a promotion board.
36 out of 37
First order of business: prove to the search engines that I own what I own. Domain verification, DNS records, the usual ritual where you paste a TXT record and Google decides whether you're worthy.
36 of 37 verified. The 37th cannot be verified, because the 37th has no DNS at all. You cannot prove you own a site that does not technically exist. There is something almost philosophical about that. I own the name. The name points at nothing. Google, reasonably, declines to have an opinion about the void. Fair enough. Even the Marine Corps didn't make you stand a post at a building that wasn't built yet.
The one rule
The whole pipeline runs on one principle: never tell a search engine anything you haven't verified yourself.
So the sitemaps are not hand-maintained lists. Hand-maintained lists drift. Every hand-maintained list in the history of computing has drifted, usually within the week, usually while its author was saying "I'll keep it updated." A sitemap full of 404s isn't neutral, either — it's a quality signal against you. You're literally handing Google a signed document that says "here are my pages" and half of them are lies.
My sitemaps are generated by crawling my own sites. Every URL that goes in the map is a URL the crawler just fetched and got a 200 from. If the page doesn't answer, it doesn't get listed. The map cannot drift, because the map is regenerated from reality every night.
And the crawler proved the point on day one. It crawled my GTA VI fan site — gta6vc.us, a labor of love I'll defend to anyone — and reported a /news page.
I did not build a /news page.
One of my own AI sessions built a /news page. On my domain. At some point. And logged it somewhere I clearly did not read closely enough. My sitemap generator, four minutes into its life, knew my estate better than I did. I am tech's bitch, and now I have the receipts in XML format.
Load-bearing comments
The nightly cron does two things: regenerate the sitemaps, then submit them. In that order. The comment in the script says:
that order is load-bearing
Because if you submit first and generate second, you're announcing yesterday's map every single night, forever, with the discipline of a machine. Automation doesn't make mistakes occasionally. It makes the same mistake at 4 a.m. daily until someone notices. That comment is there for the next AI session that decides to "clean up" the script. I know my troops.
Each run pushes roughly 297 URLs to IndexNow — the instant-indexing protocol. Which is where the research delivered a sobering finding: IndexNow reaches Bing and Yandex only. Google never adopted it. And Google's own Indexing API? Officially limited to job postings and live-broadcast events. That's it. That's the list.
Now — funny thing. One page on the fan site does carry broadcast-event schema, for the premiere everyone's counting down to. But it's a fan page about someone else's broadcast. I am not the broadcaster. Rockstar is the broadcaster. So it does not qualify, and I wrote the rule down where the whole org can see it: do not contort schema to unlock an API. The robots remember. You might fool the endpoint for a month. You will not fool the quality systems, and the penalty box has no appeals process.
The gotchas, because there are always gotchas
Two subdomains, one docroot. Same app serving two hostnames. Which means a file-based robots.txt cannot differ between them — and one site's Disallow: / was silently blocking its completely innocent sibling. One robots file, two sites, one of them collateral damage. Robots.txt is now served per-vhost from nginx, so each hostname gets its own answer. The file on disk lied by omission; the web server tells the truth per-site.
Query strings needed a rule of engagement. One parameter is a page — ?h=the-icons is a real destination and belongs in the map. Two parameters is a filter state — nobody needs ?k=cart&a=bronx indexed. I know this rule is correct because I found both failure modes personally: blanket-dropping query strings cost one site 40 real pages, and keeping everything produced invalid XML because a raw ampersand walked into a sitemap like it owned the place.
A countdown with no null guard. The fan site's premiere countdown threw an error on every page that didn't have the clock — and took the analytics beacon down with it as collateral. So the pages without a countdown were also the pages I had zero visibility into, which is a very elegant way to be blind. Guarded it. Verified zero JS errors across all six pages, because "I fixed it" without verification is just a mood.
Honest expectations
Here's the part where I resist the urge to lie to myself, which is the hardest DevOps discipline there is.
A fan site is not going to outrank the giants for the big keyword. It's not close. Its lane is the long tail — the weirdly specific questions real humans type at midnight, the "what time exactly" and "how do I" queries the big outlets don't bother answering cleanly.
And on premiere day? The traffic won't come from Google at all. It'll come from Reddit, because that's where the faithful gather, and no sitemap on Earth changes that.
But the robots can read all 37 — well, 36 — of my domains now. Every URL verified. Every map regenerated nightly from ground truth. My estate finally tells the truth about itself, which is more than I could say for its owner, who didn't know he had a news page.
Semper Fi. Submit after generating. That order is load-bearing.