Random Learning
← The journal

August 18, 2026

3 things I learned

last30days v3.3.2 · synced 2026-08-18

What I learned:

The best datasets in the corpus say data centers have been pushing retail electricity prices down, which is the opposite of what everyone is arguing about - Fortune reports new research finding that from 2015 to 2024, average retail electricity prices decreased by 3.5% for every doubling of data center capacity. A widely-shared X thread citing Lawrence Berkeley National Lab data makes the same case harder: real retail prices fell over 1¢/kWh, 7–8%, in the highest-growth states between 2019 and 2025, and the conclusion drawn is that "data centers do not inherently raise consumer electricity bills - bad utility contracts do." A second X post points at an Electric Power Research Institute study finding data centers put downward pressure on average prices through 2024. Every one of those numbers ends in 2024 or 2025. The complaint is about 2026, and that gap is why both sides can quote real research at each other.

The mechanism that actually decides your bill is not demand, it is who pays for the dedicated infrastructure - Forbes frames it as rules rather than megawatts: a tariff approved in July 2025 for projects of at least 25 megawatts "makes large data centers stand behind their own demand: pay for most of the capacity you reserve even if you use less, commit for the load-ramp period plus at least eight years." The default it replaces is the problem — techjournal.org describes the status quo as existing utility rules that "spread infrastructure costs across all customers." So the question "do data centers raise my bill" has no general answer; it resolves entirely to which tariff your utility is on.

Virginia is the live test case, and it changed the rule this month - The single biggest on-topic community thread in the window is r/Virginia at 1,137 points on Gov. Spanberger announcing that data centers must pay for new transmission infrastructure, echoed by Hacker News and a second r/Virginia thread at 963 points whose headline claims 76% price hikes and says the state now requires firms to pay for all dedicated upstream electrical infrastructure, with the governor saying it will save residents hundreds of millions. Forbes confirms the direction more soberly: several states are moving to large-load tariffs, and Virginia "recently strengthened rules requiring many new facilities to pay for dedicated ups[tream infrastructure]." Treat the 76% as a headline number from a community post, not a regulator's figure — the underlying policy change is the part every source agrees on. Worth flagging: the same 76% circulated two months ago as PJM's wholesale capacity price spike, which is a different quantity from a household's retail rate, and appears to have been re-attributed to Virginia retail bills somewhere along the way.

The politics have already detached from the economics - Whatever the price studies say, the three highest-engagement threads in the entire corpus are about resistance, not rates: 37,036 points on a high school teacher arrested for clapping in support of anti-data center activists, 27,019 points on Salem bringing a guillotine to oppose a new data center, and 23,681 points on how many ordinary people are being arrested at these hearings. Politico reads the polling as "a growing political reckoning," and it is landing: New York enacted a statewide moratorium on large data center permits in July 2026 per techjournal.org, and Newsweek covers a proposed 1-cent-per-kilowatt-hour excise tax on electricity used by data centers over 1 megawatt, with revenue split among housing, conservation, cleanup and transportation.

The grid operator's answer is to put data centers first in line to be cut off - The most concrete structural move in the window is r/technology at 1,335 points: America's largest grid wants to shed new data centers first during shortages, with 50MW-plus facilities required to bring their own generation to avoid shutoffs. That reframes the entire cost fight — instead of arguing about who pays for capacity, it makes large loads interruptible by default, which is a much cheaper way to protect ratepayers than building for a peak that may never arrive.

A quiet Bloomberg line undercuts the premise of the whole argument - Most power sought for US data centers will never materialize, which pairs with CNN noting that Americans are rallying against data centers and "surprisingly few are getting built." Utilities plan against interconnection queues stuffed with speculative requests, and if most of that demand is phantom, ratepayers can end up funding infrastructure for buildings that never open — a failure mode that looks identical on your bill to the one everyone is protesting, but has the opposite cause. r/DataHoarder is asking the same question from the other end at 418 points and 341 comments.

What is not in the corpus is a single trustworthy number for the bill impact - The most-engaged claim on X, roughly $23 billion added to consumer electricity bills, comes from a 57-like post citing "one analysis" with no named source, and the engine flagged the surrounding cluster as thin evidence. The corpus has three Forbes-class explainers on the rules, one grid-operator policy, one moratorium, one proposed excise tax, and several county-level anecdotes like Ashburn's diesel generators firing up at 1,971 points — and nothing resembling an audited national figure. Also worth holding: r/Economics at 18,161 points notes rates are up 18% against a promise to halve them, which is the actual lived number people are attributing to data centers whether or not the attribution holds.

KEY PATTERNS from the research: 1. The strongest price data runs the other way — retail prices fell 3.5% per doubling of data center capacity, 2015–2024, per Fortune — but it all predates the period people are complaining about. 2. The variable that decides your bill is the tariff, not the load: reserve-what-you-pay-for contracts at 25MW+ versus the default of socializing infrastructure cost, per Forbes. 3. Virginia just moved from the default to the contract, and it is the most-discussed policy change of the window - per r/Virginia. 4. Engagement is overwhelmingly about arrests, guillotines and moratoriums rather than rate design — the politics are running ahead of the economics. 5. The grid operator's fix is interruptibility: cut 50MW-plus loads first, make them bring their own generation - per r/technology. 6. If most requested power never materializes, per Bloomberg, the ratepayer risk is paying for capacity nobody ever uses. 7. No source in the corpus offers an audited national bill-impact figure; the viral $23 billion is an unsourced claim in a low-engagement post.

last30days v3.3.2 · synced 2026-08-18

What I learned:

The archive that nearly died this month was not killed by decaying media — it was killed by custody - Nine PBS lost access to more than 50 terabytes covering 70 years of regional television because its contracted storage vendor, Open Source Storage, simply went out of business, per Tom's Hardware and Fstoppers, which calls it "a cautionary tale on backup strategy." The bits were fine. The hardware was fine, sitting in an Iron Mountain facility. What failed was the chain of who had the legal right and the physical ability to reach them. The story broke to r/Archivists mid-window at 757 points and, as of today, r/DataHoarder reports a judge has cleared the station to retrieve the data, ruling that it owns the 50TB held on the defunct host's servers.

The preservation profession's actual position is that a backup is not preservation, and they say it in almost those words - The sharpest line in the corpus is from a university library guide, not a tech blog: "Disaster recovery strategies and backup systems are not sufficient to ensure survival and access to authentic digital resources over time. A backup is a short-term data recovery solution following loss or corruption and is fundamentally different" from preservation, per Florida Virtual Campus. Gizmodo puts the same idea as a process rather than a product: "durability comes from constant stewardship. Data is never simply stored once and for all; it is continuously recreated through copying, verification" — an archive is a verb, and the moment nobody is performing it, the countdown starts regardless of what it is written on.

Meanwhile the hobbyist layer spent the month measuring exactly the wrong variable, extremely well - The most-engaged threads in the window are all media-endurance: three years of microSD endurance testing at 4,193 points and 329 comments, four 8TB drives approaching ten years of service at 5,656 points, and a peer-reviewed study of 443,000 Backblaze drives ranking HGST most reliable and Toshiba least. This is genuinely good empirical work on which media survive. It is also the variable that did not fail at Nine PBS, and Wikipedia's lost media entry is blunt that no medium is the answer anyway — flash drives, SD cards, SSDs and hard disks "all naturally degrade over time."

The failure people actually post about is the one with no hardware in it at all - r/selfhosted at 1,066 points and 501 comments is titled simply "Well, today I lost all my data," and the corpus's most useful recurring thread is the 3-2-1 backup question — "where is your '1'?" — 241 comments of people discovering their offsite copy is a single account at a single provider they do not control. That is the Nine PBS failure mode in miniature: not rot, not a crash, but one dependency nobody stress-tested because it was a company rather than a component.

What professionals keep coming back to is owning the substrate, and tape is still what that looks like - The counterexample running quietly through r/DataHoarder is the IBM TS3500 tape library thread at 429 points — a physical robot in a room you have keys to. The tape story in the window is not romantic, either: one operator reported on X that new Ultrium LTO-10 drives "simply do not work," having gone through two units across multiple workstations that could not write a single file larger than 1GB. Even the medium built for archival custody has a supply chain, and the supply chain has a bad month sometimes.

The volunteers are load-bearing infrastructure, and they are visibly precarious - An army of amateur archivists racing to save US history in national parks drew 1,222 points; the same community spent the window on DVDBeaver at risk of shutting down at 500 points, and r/ReelToReel surfaced a Kazakh collector preserving Soviet tape recorders that defied the musical iron curtain. None of these have institutional funding, and all of them are single points of failure in exactly the way Nine PBS's vendor was.

And there is a newer way to lose an archive that no storage strategy addresses - r/Archivists documents automated moderation erasing a decade of r/AskHistorians work — the data was never on your media, never in your custody, and was deleted by the platform's own tooling. For anything that lives only inside someone else's service, the medium question and the vendor question collapse into the same question, and you do not get to answer either one.

Honest note on the evidence - This brief's live layer is real but narrow. Reddit's public search endpoint returned 403 and 429 across passes, so the community evidence comes from subreddit listing discovery rather than topical search, and two earlier framings of this topic — analog tape preservation, then long-term tape archiving specifically — returned corpora dominated by evergreen hobbyist video and unrelated cloud-computing threads. Several threads cited above surfaced in an earlier pass on this same slot; the footer below reports the final pass only. The Nine PBS thread, the endurance tests and the library guidance are all within the window and independently corroborated across Tom's Hardware, Fstoppers and the public-media trade press.

KEY PATTERNS from the research: 1. The month's marquee near-loss was a custody failure, not a media failure — a vendor went under and took access with it, per Tom's Hardware. 2. Archivists draw a hard line the hobbyist world usually blurs: backup is short-term recovery, preservation is continuous stewardship - per FLVC. 3. Durability is a process of continuous copying and verification, not a property of a disk - per Gizmodo. 4. The highest-engagement community work this window measured media endurance — good data, wrong failure mode. 5. "Where is your 1?" is the load-bearing question in 3-2-1, and for most people the honest answer is a single vendor account - per r/selfhosted. 6. Owning the substrate still means tape for serious archives, and even LTO had a hardware complaint in the window. 7. Volunteer archivists are doing institutional work with no institutional backing, and at least one long-running film archive is currently at risk of shutting down. 8. Platform-side automated deletion is a loss mode that no personal storage strategy can reach - per r/Archivists.

last30days v3.3.2 · synced 2026-08-18

What I learned:

The traffic is real, it is most of the web now, and your analytics cannot see any of it - The number repeated across the corpus is that bots account for more than half of all web traffic — LeadJaw puts it at 57.5%, and SitePro News frames the consequence as your site having two audiences with only one showing up in the dashboard, because "AI crawlers are not polite browsers." The most concrete first-person measurement in the window is an operator who pulled a client's server logs: GPTBot 400+ visits a day, Meta-ExternalAgent 220, ClaudeBot 180, PerplexityBot 90 — 890 AI visits daily, and "his Google Analytics showed zero of them." The mechanism is boring and total: "GA won't show bots at all... it runs on javascript so crawlers never fire it. server logs are the real answer but parsing them sucks."

The institutions getting hit hardest are the ones holding public collections, and they are saying so in conference talks rather than blog posts - The two on-topic videos in the entire window are both Drupal community sessions about exactly this. AI Crawlers Are Crushing Your Website, from the Drupal Association, describes "a new kind of bot... scraping public government websites to feed large language models" that "aren't malicious, but they hit fast, often ignore robots.txt, and can over[whelm]" the server. Surviving the Swarm is blunter about the state of the art: "yesterday's strategies and defenses are no longer enough." That is the tell for where this problem actually lives — under-resourced public sites with large document collections and no CDN budget.

The defense that won the year is a proof-of-work gate with a cartoon jackal on it - Anubis sits at 21,000 stars in Go with 357 open issues and a one-line pitch — "weighs the soul of incoming HTTP requests to stop AI crawlers" — shipping through v1.27.0 in the window. Its companion is the blocklist approach: ai.robots.txt at 4,100 stars, "a list of AI-related crawlers of all types, regardless of purpose." The pairing is the whole strategy in miniature: a list for the crawlers that obey rules, and a computational toll for the ones that don't.

Blocking is only one of three answers, and the other two are stranger - The second is to serve bots something different: TIME is serving AI bots a different website, with ads built in, which drew 267 points and 110 comments on Hacker News — a publisher deciding that if the crawler is coming anyway, it should be monetized rather than refused. The third is to give up on refusal entirely and optimize for the machine reader, which is the premise of the GEO crawl-to-refer analysis and the reason a consultant's viral finding lands the way it does: disable JavaScript on a typical product page and "price gone. reviews gone. product specs gone" — which is what GPTBot, ClaudeBot and PerplexityBot see when they arrive.

Not all crawler traffic is the same thing, and the split matters for what you should do - New Market Pitch reports that broad crawlers produce about 85% of identifiable AI requests while fetchers tied to a live user question produce 15%, with fully delegated agents smaller still. HUMAN Security draws the same line qualitatively, describing well-behaved scraping that "respects robots.txt directives, maintains reasonable request rates." Blanket blocking treats the 15% that arrived because a human asked a question the same as the 85% harvesting for training — which is why the blocklists and the pay-per-access experiments keep diverging.

The load complaint is the small version of the argument; the memory complaint is the big one - The highest-engagement item anywhere near this topic is As AI eats the web, the internet's collective memory is disappearing at 937 points and 987 comments, paired with AI;DR (AI; Didn't Read) at 951 points and 579 comments. The worry there is not server bills — it is that the crawl-to-referral bargain has broken, so the sites being read are no longer being visited, and the incentive to keep publishing the record erodes. Server load is what makes an archivist install Anubis; the referral collapse is what makes them wonder why they are hosting at all.

And there is a failure mode where the archive is destroyed by AI without a crawler involved - r/Archivists documents automated moderation erasing a decade of r/AskHistorians work. No scraper, no bandwidth, no robots.txt — the platform's own tooling removed the collection. For anything hosted on someone else's service, the defenses in this brief do not apply at all.

Honest note on the corpus - The community layer here is broad but noisy: the subreddit sweep across r/selfhosted, r/sysadmin and r/webdev returned mostly unrelated high-engagement threads, so the load-bearing evidence is Hacker News, two Drupal conference talks, two GitHub projects, one operator's server logs on X, and a set of vendor and SEO analyses whose framing is commercial. Notably absent from the window: any first-party report from a large library or national archive putting numbers on their own bot traffic, and any concrete data on what per-crawl payment schemes are actually earning.

KEY PATTERNS from the research: 1. Bots are now the majority of web traffic — 57.5% per LeadJaw — and JavaScript-based analytics structurally cannot see them. 2. Server logs, not the analytics dashboard, are the only honest measurement; one site logged 890 AI visits a day that GA reported as zero - per this operator. 3. Public-sector and collection-holding sites are the ones describing real outages, and they say crawlers routinely ignore robots.txt - per the Drupal Association. 4. The defensive stack that consolidated this year is a proof-of-work gate (Anubis, 21K stars) plus a shared blocklist (ai.robots.txt, 4.1K stars). 5. Three strategies are live at once: block them, serve them a different site with ads, or optimize for them — and publishers are picking different ones. 6. About 85% of AI requests are broad training crawls versus 15% user-triggered fetches, so blanket blocking cuts both - per New Market Pitch. 7. The argument with the most energy behind it is not bandwidth but the disappearance of the web's collective memory - per The Walrus at 937 points. 8. Platform-side AI moderation can erase an archive outright, which no crawler defense addresses - per r/Archivists.

Provenance — 2026-08-18

Redacted by design: this records the funnel shape, not the private source links or personal capture notes. Raw self URLs and why? text are never written here.

Source entries (3 picked from a pool of 5)

A thin pool today — five eligible entries, all captured in the last 48 hours. Three were picked, weighting each entry's own capture note heaviest and then spreading across fields.

  • A saved news item on young people's distrust of AI executives (tags: technology, ai, youth, trust, ceo, surveys), captured 16 August. The capture note is the most emotionally loaded of the pool and names two specific grievances rather than a general mood — jobs and community. The fan followed the community half, because that is where the last 30 days have hard numbers and the jobs half is better covered elsewhere in the index.
  • A saved magazine feature on an early magnetic tape recorder and its downstream effect on broadcasting (tags: recording, technology, radio, history), captured 17 August. The capture note registers curiosity about the past with no present-day hook, so the fan went forward in time rather than deeper into the artifact: what holds recorded material now, and what destroys it.
  • A saved visual archive of a computing magazine's back catalogue (tags: visual, archive, media), captured 17 August. A delight note about the artifact; the fan went to the conditions such archives currently survive under.

Two entries were left in the pool: one strategy essay whose capture note was the weakest of the five ("good to know"), and one developer-tooling gallery that would have made a third consecutive AI/agent-tooling day in an index already heavy with them.

Domain spread: energy and utility regulation · digital preservation · web infrastructure.

The 12 adjacent candidates

From the AI-executive-distrust item: 1. Whether entry-level hiring actually collapsed in the jobs AI was supposed to take 2. What actually happens to a town's power bill when the data center arrives ← picked 3. How tech-executive credibility is polled and what moved it this year 4. Whether announced AI layoffs match the actual filings

From the tape-recorder history: 5. How the laugh track was actually built and why it died 6. What analog tape does that digital doesn't, per people still recording to it 7. The magnetic tape deadline: getting audio off decaying reels before it's unplayable ← picked, then reframed (see below) 8. What tape-emulation plugins get wrong about saturation

From the magazine archive: 9. How people put whole magazine runs online without getting taken down 10. Where the Internet Archive's legal position stands after the publisher suits 11. How much of the print and early-web design record is already lost 12. What archives are doing about AI crawlers hammering their servers ← picked

Near-dup guard: 0 of 12 flagged against a 177-topic index.

Narrowing to 3

One topic per source entry, three different fields. #2 over #1/#3/#4 because it is the concrete, locally-measurable version of the capture note's "community" grievance and the window is full of primary policy movement. #7 over #5/#6/#8 as the only candidate in that cluster with a real stake rather than a hobby interest. #12 over #9/#10/#11 because the legal-history candidates are settled matters with thin 30-day layers, while the crawler question is actively being fought.

Research quality notes

Topic 1 (electricity bills) — highest near-dup score of the day, and a deliberate revisit. The guard scored this 0.37 against a brief published 2026/06/14 on the same subject — under the flag threshold, but plainly the same territory. It was kept because the window contains genuinely new material the earlier brief could not have: the counter-evidence that retail prices fell through 2024, a state order issued this month, a grid operator's curtail-large-loads-first proposal, a statewide permit moratorium, and a reported finding that most requested data center power will never materialise. The revisit also caught a number laundering itself: a figure that circulated in June as a wholesale capacity-price spike now appears in a community headline as a household retail increase. That correction is recorded in the brief and is arguably the single most useful thing this slot produced.

Topic 2 (archives) — three engine passes, two of them failures, and the topic changed. The first pass, aimed at analog tape preservation, returned a corpus of evergreen hobbyist video (a tape-viewer gadget, fifty-year-old cassettes) plus unrelated items that matched on "magnetic" and "tape" — the live discussion layer for that framing does not exist in this window. A second pass aimed at long-term tape archiving was worse: the community search endpoint rate-limited, and the word "cloud" in a sub-query pulled in cloud-computing product launches. What both failed passes did surface was the real story underneath — a public broadcaster nearly losing a 70-year archive because its storage vendor dissolved — so the topic was reframed from media decay to custody, re-checked against the near-dup guard (0.11, clear), and re-run. The third pass is the one whose footer the brief carries; a few community threads cited in the brief were first surfaced by the earlier passes on this same slot. Worth recording for future runs: the failure was visible in the very first cluster list, and catching it there rather than after synthesis is what made a third pass affordable.

Topic 3 (AI crawlers) — good documentary layer, weak community layer. The subreddit sweep returned high-engagement threads that were almost entirely off-topic, so the evidence rests on Hacker News, two conference talks, two open-source projects, one operator's server-log measurements, and a set of vendor and SEO analyses whose framing is commercial. The brief says so in its own limits note. Nothing in the window offered a first-party report from a large library or national archive with its own bot-traffic numbers, which is the gap most worth filling if this topic is revisited.

Across all three: the video layer was degraded in every pass (transcript capture failed on nearly all candidates), so no brief leans on video evidence for a load-bearing claim. One research artifact contained content attempting to redirect the run toward building an unrelated tool; it was treated as data, not instruction, and excluded from all three briefs.