Random Learning
← The journal

August 24, 2026

3 things I learned

last30days v3.3.2 · synced 2026-08-24

What I learned:

Blocking yourself from the data is free, takes a day, and lasts two years - The thing importers actually do about being visible in bill-of-lading records is not a lawsuit or a shell company, it is a form. Under 19 CFR 103.31(d) an importer or consignee can ask CBP for confidential treatment of its own name and address and of its shippers' names and addresses, across 22 manifest data elements including consignee and notify party. There is no fee, the electronic tool can process a request in under 24 hours, the grant covers every U.S. port of entry, and it runs two years before it has to be renewed. The catch is string matching: the name you submit has to match exactly what your carrier or filer typed into the Automated Commercial Environment, so a company that files under four spellings has to enumerate all four or it leaks through the one it forgot.

The records return a freight forwarder unless you know which bill you are reading - The sharpest technical point in the whole window came from a trade-data vendor blog rather than a forum. EximAgent lays out the two failure modes plainly: "A company's absence is not proof they do not import," and when a forwarder consolidates cargo the master bill names the forwarder while "the house bill underneath it carries the real consignee." Their line for people who skip that distinction is that you "will build a prospect list of logistics companies." They also flag that declared values are filed for customs purposes and are not usable as margin inputs, which quietly kills the most common thing people want the data for.

The live 30-day discussion layer on this is close to silent - This is the honest finding. The engine pulled 46 items and almost every one came back entity-miss demoted. The sourcing and logistics communities that should be having this conversation were having a different one: r/supplychain spent the month on salary comparisons and specialty choice, r/logistics on war stories and pallet moves. Manifest mining as a tactic did not surface as a thread anywhere in the window. Nobody is arguing about this in public right now, which is itself the signal: it is settled practice on both sides, not a live controversy.

Supplier opacity is being fought upstream of customs, at the paperwork layer - Where importers did complain about being unable to see through a supplier, it had nothing to do with manifests. On r/ecommerce a brand bringing a custom plush toy to market hit a wall registering a law-label URN because the Chinese manufacturer "flat out refused to share their local tax ID / business registration code (Unified Social Credit Code), CEO details" or anything else identifying. The same subreddit ran a full factory-cluster teardown pricing a MagSafe charger at $3.60 FOB against $44.99 retail, an 82% spread, which is exactly the reverse-engineering that manifest data is supposed to enable and which this poster did without it.

What put CBP data in front of a technical audience this month was misuse, not sourcing - The only CBP story with real engagement on Hacker News was about officers and contractors abusing government databases to look up exes and colleagues, carried by both Wired and Techdirt. That is a different database than the public manifest file, but it lands on the same nerve: who gets to query trade and traveler records and with what accountability. Meanwhile CBP's own enforcement channel goes the other direction on disclosure, telling tipsters via e-Allegations that once a case closes it "is unable to disclose specific enforcement actions, if any, due to restrictions imposed by the Trade Secrets Act, the Privacy Act, and CBP regulations."

The vendor layer is stable, old, and unglamorous - ImportGenius is still selling U.S. import and export records "at the bill of lading level" and leaning on being nearly two decades old; Panjiva sits inside S&P Global; Descartes Datamyne rounds out the named trio in practitioner writeups like EcomCrew's. All of them are reselling the same government pipe, which CBP describes as its "electronic Single Window platform for all trade processing, including all Manifest, Cargo Release, Post-release, Export" data on the ACE page. The only forward-looking take in the social layer was @mattklucas arguing that AI document readers for manifests are an on-ramp rather than the product, and even CBP's own field accounts frame manifests as a verification artifact, with @DFOFlorida noting that when an x-ray anomaly appears officers check "that it matches the manifest."

KEY PATTERNS from the research: 1. The counter-move to being discovered is administrative, not legal, and it is cheap - a free two-year confidentiality grant processed in under a day, per CBP 2. Exact-string matching against ACE filings is the real weakness in confidentiality - one unlisted name variant defeats the whole request, per CBP 3. Master bill and house bill are the difference between finding a factory and finding a forwarder, per EximAgent 4. Absence from the data proves nothing, and declared values are not margin data, so the two most tempting inferences are both unsafe, per EximAgent 5. Practitioners are hitting supplier opacity at the registration-document layer, well before manifests are relevant, per r/ecommerce 6. The live community conversation on manifest mining is effectively dead this window, which reads as settled tradecraft rather than lost interest, per r/supplychain

last30days v3.3.2 · synced 2026-08-24

What I learned:

AIFS has stopped being an experiment and become a line item in the model list - ECMWF has run the AI Forecasting System alongside the physics-based IFS since February 2025, and ECMWF Open Data now ships four IFS and four AIFS cycles a day at 00, 06, 12 and 18 UTC on the same rolling archive. You can see the shift in how forecasters write: a Belgian outlook from @ltrullem lists its sources as "ICON-AIFS-ECMWF-GFS" with no ceremony, @gpvweather posted IFS and AIFS tracks side by side for Typhoon 18 off Japan, and @allindiaweather ran a straight AIFS 72-hour rainfall call for Odisha and East UP. On Hacker News, the answer to "where can I actually see these" was flat: "AIFS directly by ECMWF and AIGEFS by NOAA. Every vibecoded weather app these days has them."

The 2026 season produced the first public case of a hurricane modeler openly deferring to the machine - @AndyHazelton posted on 22 August that on a Main Development Region wave, GFS and ECMWF were both aggressive on development while "AIFS and Google DeepMind are a lot less excited about development and hold off any chances until much later," and then landed on the line that matters: "At this point I think I'd defer to the AI models a little more after seeing their performance the last year or so, but we'll see." Three days earlier the same account sided against the GFS on which wave would develop, on the grounds that "Euro, AIFS, and FNV3" agreed with each other and "the GFS usual biases in this region" did not. That is a working forecaster using three machine-learning models as a consensus block against a physics model, in public, mid-season.

The sharpest practitioner complaint is not that the AI is wrong, it is that the AI answer is climatology wearing a lab coat - @RyanWeather read out the 12z DeepMind and ECMWF AIFS-ENS update on 23 August, called a good chance of a wave becoming a hurricane in 7-10 days with "under 10% chance impact to US mainland," and then killed it in one line: "That's about climatology or 'going rate' for any Cape Verde style 'cane. Not really adding value." @Weatherunited1 hit the same wall from the other direction, noting that ECMWF and ECMWF-AIFS had both downtrended over four runs while the GFS spun up a hurricane, and concluding "at this point, ensembles are also a bit useless more than 5 days out."

The measured gains are real, and narrower than the press release - ECMWF's own AIFS work puts deterministic AIFS 12 to 24 hours ahead of the IFS on anomaly correlation with roughly 10% improvement through the troposphere, and AIFS-CRPS 5 to 20% better than the IFS ensemble on CRPS and RMSE for upper-air fields, while explicitly degrading at 100 hPa and above; the same update reports about 30% 2 m temperature RMSE degradation in the Arctic scored against analysis for January to March 2026. On the cyclone side, DeepMind describes WeatherNext Cyclones as iteratively predicting both global weather patterns and fine-scale cyclone tracks out to 15 days, and claims it sees about one day further ahead than leading operational models — which is precisely the claim Hacker News went after, below. WIRED anchors it on Melissa: five days before landfall, 80% confidence of a Category 5 strike on Jamaica.

Hacker News immediately corrected the "extra day of warning" framing, and the correction came from inside the field - the thread hit 449 points and 130 comments, and the most load-bearing comment is @counters: "with modern forecasting tools, we anticipate tropical cyclones to develop 5-10 days before they ever threaten landfall," so the improvement is "better interpreted as a modest reduction in forecast uncertainty - the 'cone' on the hurricane track map gets a little narrower," and "nothing actually changes on-the-ground for really any consumer of hurricane forecast data anywhere in the world." @TaupeRanger had already restated the claim precisely: "the model can reach a given level of forecast accuracy roughly a day farther in advance." Both explicitly said this is not a knock on DeepMind, it is a statement about how good forecasting already was.

Where the machine-learning forecasts still fall down has a shape, and it is the tails - the Science Advances result summarised by Carbon Brief and Physics World is that AI models systematically err toward normality, under-forecasting record heat and over-forecasting record cold, with the error growing in proportion to the margin by which the record is broken. arXiv work on sudden-turning typhoons traces the track failures to an inability to resolve fine TC structure, and a July preprint names the ceiling directly: high-resolution AI forecasting is constrained because decades of reanalysis only exist at 0.25 degrees. There is also a mundane operational gap - @MassachusettsWx points out that GFS gives hourly output to 120 hours and ECMWF to 90, while the AI models step at 6 hours, and that HourGlass exists purely to fill in the blanks until the AIFS/ENS v3 upgrade later this year.

The dependency nobody in the window disputes is that AI models eat physics output for breakfast - @sunshinesnacks put it plainly: "pretty much all of the AI weather prediction models are trained on ECMWF ERA5, which is kinda like a numerical weather prediction model run to forecast at t=0," and @tcumulus extended it to runtime - "AI weather models depend greatly on the NWP/physics used in for reanalysis and initial conditions." The escape hatch is already named: @micro2588 flagged AIFS-DOP, ECMWF's experimental direct-observation model that has become "competitive with their physics based IFS model on certain metrics just in the past year." Two other comments cut against easy triumphalism from opposite sides - @jeffbee noting that physics models "rely heavily on humans looking at the output and the evaluation of the output to discard wacky runs," and @ghm2199 quoting GraphCast's own admission that an MSE objective "encourages it to express its uncertainty by spatially blurring its predictions," then asking whether you would issue an evacuation order 30 miles from a hurricane centre on an uncertainty estimate the model cannot explain.

The biggest weather thread on Reddit this month was not about AI at all, it was about the observations underneath it - "DOGE destroyed how America forecasts its weather" pulled 578 points on r/weather and 248 more on r/meteorology with the subtitle doing the work: fewer weather balloons and wave-monitoring buoys launched to collect the data forecasts are built from. @alternator drew the connecting line on HN, arguing that headlines about neural nets "outperforming NOAA" helped justify the cuts while "the industrial models utterly rely on government data for inputs." Meanwhile NOAA shipped its own AI stack - AIGFS, AIGEFS and HGEFS went operational on 17 December 2025, AIGFS running on up to 99.7% less compute and AIGEFS extending skill 18 to 24 hours past GEFS, and 2026 is the first season in the National Hurricane Center's 70-year history that AI guidance entered operational hurricane forecasts.

The code layer consolidated in exactly this window - google-deepmind/weathernext landed on 6 August under Apache 2.0 carrying WeatherNext 2 plus the older GraphCast code, and is already at 7.6K stars with 77 open issues, against 1.0K stars for microsoft/aurora and 1.1K for NVIDIA/earth2studio. The quiet casualty is ecmwf-lab/ai-models, the 585-star wrapper that was how most people ran GraphCast and Aurora on ECMWF data - its README now leads with an archived project-maturity badge. And for anyone tracking live guidance rather than running it, the HN thread's recommendation was Zoom Earth and Tropical Tidbits, with @jen729w pasting raw JTWC-style discussion text on Typhoon Dolphin's "trochoidal Z motion" as an example of what is now routinely public.

KEY PATTERNS from the research: 1. Working forecasters have moved from evaluating AI models to voting with them - the AI models now function as a consensus block that a physics model can be argued against, per @AndyHazelton 2. The failure mode practitioners name is uselessness, not error - an AI ensemble that returns climatology is "not really adding value," per @RyanWeather 3. The "extra day of warning" number is a narrower cone, not extra preparation time, and the correction came from inside the field, per @counters on HN 4. Every headline gain has a stated regression next to it in the same paper - 5 to 20% better ensemble scores with degradation above 100 hPa, per ECMWF's AIFS-CRPS paper 5. AI models still fail in the tails by regressing toward normal, and the error scales with how badly the record is broken, per Carbon Brief 6. The live argument is not about models at all, it is about whether the observation network that feeds them is being defunded on the strength of their own benchmark wins, per r/weather

last30days v3.3.2 · synced 2026-08-24

What I learned:

The deletion clock actually started, and it is a recurring one - As of 1 August 2026 every registered data broker must pull the DROP list and delete, then do it again every 45 days, per Privacy Rights Clearinghouse. @CalPrivacy framed it the night before as "a major milestone in California privacy... It's simple. It's free. It's your right." TrustArc counts more than 600 registered brokers now in scope, and the Governor's office confirmed the obligation started on the 1st. Over 300,000 Californians have enrolled since the consumer side opened in January, and deletion violations carry $200 per consumer per day, per Privacy Rights Clearinghouse.

The first enforcement actions were about registration and opt-out friction, not about deletion - @CalPrivacy announced its first action under both the CCPA and the Delete Act on 11 August: "LocateSmarter LLC must pay $116k and change its practices after failing to register and follow data minimization requirements." The specific sin was making people hand over sensitive data to opt out - Verified Privacy VPN put it bluntly in a video titled "California Fines Data Broker That Required Your SSN to Opt Out," noting the form asked for the last four digits of a Social Security number. Fried Frank adds that LocateSmarter also has to pay the $6,000 DROP registration fee. Two days later @CalPrivacy hit Cybba, Inc. for $52k: "In its second decision in less than a week... With a pattern of robust enforcement, CalPrivacy's message is clear: data brokers must follow the law." Hunton puts the exact Cybba figure at $52,400, and Fisher Phillips dates both settlements to 10 August.

The repopulation question has a statutory answer, and it is the most interesting design choice in the law - DROP is not a one-shot delete; it is a standing suppression list. Privacy Rights Clearinghouse states the rule plainly: if a broker has re-acquired your information, they delete it again on the next 45-day pass, and the Delete Act separately bars brokers from selling or sharing newly collected information about anyone with an active deletion request. Brokers who decline a request have to report the denial count to CalPrivacy and cite the specific statutory exemption they are relying on. That is the mechanism the DIY opt-out world never had - manual removals have always been point-in-time, which is exactly why records came back.

The pre-DROP measurement of whether removals stick is grim, and it is the number to beat - Consumer Reports ran a four-month study with Tall Poppy across seven services (Confidently, DeleteMe, EasyOptOuts, IDX, Kanary, Optery, ReputationDefender) against 13 prominent people-search sites: of 332 identified data instances for 28 participants, only 117 - 35 percent - came down within four months. Doing it yourself beat every paid service, with manual opt-outs clearing 70 percent within one week. EasyOptOuts was the best automated performer at 65 percent. The researchers also found financial partnerships between some removal services and the people-search sites they were removing from, which The Record picked up as the study's sharpest finding.

Even a "successful" people-search removal leaves residue - Incogni walks through the USA People Search opt-out and lands on the caveat that matters: your data should be gone from results within 72 hours, but "it'll still show up in the sponsored search results on the site, though." Clearnym tells people to save evidence of the completed removal and re-contact with the profile URL, citing a BBB complaint documenting a case where personal information kept appearing in search results after an opt-out. Scope is the other residue: DROP is a California right, and Spokeo, BeenVerified, Radaris, Acxiom, LexisNexis and CoreLogic are in scope only because they clear California's registration thresholds.

The self-hosted crowd is building around the same problem and mostly ignoring the law - yaelwrites/Big-Ass-Data-Broker-Opt-Out-List sits at 6.8K stars with a crucial/high-priority symbol system, and digisamroc/eraser (183 stars, Go) pitches itself against exactly the sites DROP covers: "You know those sites like Spokeo, BeenVerified, and Whitepages that have your home address, phone num..." while sending GDPR/CCPA requests to 750+ brokers for free. On the paid side CNBC notes EasyOptOuts is $20 a year but scans only 100+ sites where DeleteMe reviews 750+.

The loudest privacy community in the window barely mentioned any of this - r/privacy generated 35,450 upvotes across 21 threads in the last 30 days and almost none of it was about the Delete Act. The front page was Flock cameras: a grand jury declining to indict an Ohio man who destroyed one at 5,658 upvotes, Comcast turning routers into motion detectors at 2,715, and someone showing up to San Diego's city council dressed as Darth Vader to say "the emperor is a fan of Flock, and we must continue utilizing Flock technologies so that we can follow and surveil the rebel scum." The DROP-adjacent threads were quiet and contrarian: one r/privacy poster asked whether removing your old data is even beneficial, reasoning that once "previous emails, errors, and even fake data" are wiped, "every new registration gives brokers the most recent data on you."

The one piece of real explanatory content came with the regulator sitting in the chair - Techlore pulled 6,701 views and 261 likes on "How One State Just Forced Every Data Broker to Erase You (DROP Explained)," an interview with CalPrivacy Executive Director Tom Kemp covering how DROP works, what the video describes as fines of up to $40 million a day for non-compliance, and what to do if you live outside California. That last question is the one the rest of the country is stuck on: only California runs a centralized platform, while Oregon, Texas and Vermont require broker registration with no equivalent portal.

KEY PATTERNS from the research: 1. The enforcement so far punishes registration failures and hostile opt-out forms, not missed deletions - $116k against LocateSmarter for demanding SSN digits and $52,400 against Cybba for never registering, both settlements dated 10 August and announced over the following week - per @CalPrivacy 2. DROP's answer to repopulation is architectural rather than punitive: a standing suppression list re-processed every 45 days plus a ban on selling newly collected data about anyone with an active request - per Privacy Rights Clearinghouse 3. The benchmark DROP has to beat is 35 percent removal in four months from paid services, against 70 percent in one week from doing it manually - per Consumer Reports 4. Removal and invisibility are not the same thing - the same guides that confirm a 72-hour takedown also note the record persisting in sponsored results and BBB complaints about data reappearing after opt-out - per Incogni 5. Nobody in the enthusiast privacy community is treating this as the story - a 6.8K-star opt-out list and a 183-star Go tool keep growing while r/privacy spends its 35,450 upvotes on license-plate cameras - per r/privacy

Provenance — 2026-08-24

Redacted by design: this records the funnel shape, not the private source links or personal capture notes. Raw self URLs and capture-note text are never written here.

Fuel

The circuit-breaker failed on the first read — eligible_pool: 0, exit 2 — and this was the stale-cache false negative for the third consecutive day. A git fetch showed the local clone 25 commits behind the remote, which is the largest gap yet and made the diagnosis unambiguous: a genuinely exhausted pool and a badly stale clone produce the same zero. Pulling brought the pool to 11 and the breaker passed on re-run, with three days of runway. This has now recurred often enough to stop being an anomaly. The breaker reads the working tree at step 1, and the collector that updates that working tree is step 2, so the measurement is structurally taken before the update. The fix belongs in the script — either have fuel.py pull first, or have it fail distinctly when the clone is behind its remote — rather than in a daily manual workaround.

Source entries (3 picked from a pool of 11)

An unusually healthy pool, and an unusually correlated one. All 11 entries were captured in a single browsing session on 2026-08-23, and nearly all of them were public-records and discovery tools. That correlation, not scarcity, was the selection problem today.

The capture notes split cleanly into two kinds: three written in the operator's own words, and eight that were text pasted from the destination page. Per the standing weighting — the note is the heaviest signal because it is the only part the operator actually wrote — the own-words group carried the pick. Two of the final three came directly from it: one trade-data tool and one weather-visualisation tool.

The third own-words entry was deliberately passed over. Its natural fan-out — open-source replacements, de-bloating, self-hosting — collides with three already-published briefs from 17, 19 and 21 August. It was replaced with an entry from the dominant cluster, whose adjacency is unexplored. Final domain spread: international trade, atmospheric science, and internet exposure. That is a genuine range rather than a rut, which was the goal.

The 12 adjacent topics

From the trade-data entry: 1. US customs bill-of-lading data and supplier discovery → picked 2. What 2026 tariffs did to small-scale importers' sourcing decisions 3. The end of de minimis and what it changed for small ecommerce sellers 4. How importers screen suppliers under forced-labour enforcement

From the weather-visualisation entry: 5. Which weather model forecasters actually trust 6. How sailors and pilots use weather visualisation in practice 7. AI weather models in operational forecasting → picked 8. What personal weather stations contribute and where the data goes wrong

From the internet-exposure entry: 9. What internet-wide scanning finds exposed in 2026 10. The legality and ethics of mass port scanning 11. People-search data brokers and the deletion laws aimed at them → picked 12. Wardriving databases and the privacy of Wi-Fi based location

The automated near-dup guard flagged none of the twelve, and — per the standing lesson that the guard misses subject revisits — related() was also run at this step rather than at provenance time. It returned nothing for any of the twelve. That empty result is itself informative: with 189 published topics, this index has no coverage at all of trade, meteorology, or privacy law. The pool's correlation pointed somewhere genuinely new.

Narrowing to three

Curiosity and freshness pointed the same way for once, so the narrowing was cheap. The deciding criterion was the one added to this list two days ago: prefer a topic anchored on a named organisation over a topic that is a concept. Each of the three picks resolves to an institution with a dated action in the window — a customs authority and a specific regulation, a forecasting centre and its model, a state privacy regulator and a compliance deadline that fell inside the last 30 days.

Overlap was checked in both directions. #5 was dropped because it overlaps #7, and #9 was dropped in favour of #11 because #11 had a hard date inside the window.

Research notes

Three engine runs for three briefs, all landing on the first attempt. That is worth recording against the nine runs needed for three briefs on 22 August. The difference was entirely step-5 discipline: every topic string was title-shaped and anchored on a proper noun, and every seed avoided the phrasings known to trigger the engine's comparison mode. The lesson from two days ago transferred cleanly and cost nothing to apply.

Two process notes worth carrying forward.

Leak avoidance moved upstream. Two of the three picked entries are well-known enough that the engine would plausibly cite their own domains, which trips the gate's privacy check at step 10 and forces a post-hoc citation swap. Rather than wait for that, each research run was instructed at dispatch to cite equivalent public sources instead. No brief cites a picked domain, and the privacy check passed with no rework. Preventing this at step 6 is materially cheaper than repairing it at step 10.

One brief carried an unsourced claim, and the grep caught it. The weather brief asserted a five-day cyclone track-error comparison — three specific distances and a named competing model. Grepping the raw evidence found zero support for any of it: the competing model's name appears nowhere in the corpus, and neither do the figures. Every other number in the same bullet verified. The claim was cut and replaced with what the evidence actually supports. This is the second time a confident, well-formatted, entirely unbacked claim has appeared in an otherwise sound brief, which suggests it is a standing property of the synthesis step rather than bad luck. The number-by-number grep against the raw file should be treated as mandatory, not optional.

Everything else verified against the raw evidence: the confidentiality-request mechanics and its exact-match weakness, the removal-service study's participant and instance counts, both enforcement fines, the enrolment and broker-registration figures, the ensemble scoring gains and their stated regressions, and the compute-reduction figure. One internal inconsistency was corrected, where settlement dates had been conflated with announcement dates in a summary line while the body had both right.

Corpus honesty: the trade brief has effectively no live 30-day discussion layer, and says so in its own text rather than dressing up an evergreen topic as a current one.