HN Summaries - 2026-09-02

Top 8 Hacker News posts, summarized


1. Claude Fable 5.1 and Claude Mythos 5.1

HN discussion (797 points, 766 comments)

Anthropic has released Claude Fable 5.1 (generally available) and Claude Mythos 5.1 (trusted access only for cybersecurity and life sciences). Fable 5.1 reduces costs by approximately 25% for typical workloads and up to 45% for highly agentic tasks, achieved by cutting cache read pricing 75% to $0.25 per million tokens. Enterprise Frontier Safeguards (EFS) now allow customers to retain full data control on their own cloud infrastructure while maintaining misuse detection. Safeguards have been refined to produce 60% fewer false positives in cybersecurity, and Fable 5.1 can now identify software vulnerabilities for defensive purposes. Mythos 5.1 demonstrates advanced scientific capabilities: designing protein binders with 10x higher affinity and 50% hit rates, creating a high-resolution Venus elevation map from 30-year-old radar data, and optimizing GPU kernels for 2.5x speedups in computational biology. Safety evaluations show Mythos 5.1 remains below the next risk tier for chemical/biological weapons and within the lower cyber risk category. Alignment improved with less reward hacking and motivated reasoning, though gaps remain in long-context and multi-agent scenarios. Anti-distillation mechanisms have been strengthened. Two trusted access programs (Cyber Verification and Life Sciences Verification, the latter with US government partnership) gate Mythos 5.1 access. EU AI Act compliance includes invisible watermarking and a detection API. Fable 5.1 is available now across all major cloud platforms at $10/M input and $50/M output tokens.

Commenters focused heavily on pricing and accessibility. Several noted the cache read discount significantly benefits long-running agentic workflows, while others argued the base price remains too high for routine coding compared to alternatives like GPT-5 or Grok. Multiple users reported receiving usage quota resets with the release. Complaints persisted about Fable 5's over-sensitive alignment filters forcing fallback to Opus, with uncertainty whether 5.1 improves this. The absence of Haiku 5 drew repeated questions. Technical users analyzed the anti-distillation patches—blocking manual context editing for new API accounts and preventing Haiku from repeating other models' thinking blocks—as targeting specific chain-of-thought extraction techniques. Criticism targeted Anthropic's "democratization" framing given Mythos 5.1's restricted access, and several commenters described the announcement's prose as verbose, dense, and characteristic of AI-generated text. One user refused to test without a free trial period.

2. AnkiDroid: Google Play no longer allowing Open Collective donation link

HN discussion (801 points, 233 comments)

AnkiDroid, a free open-source flashcard app with over 10 million installs, faces removal from Google Play worldwide (except India and Russia) by September 11, 2026, unless it removes its Open Collective donation link. Google's Payments Policy prohibits apps from directing users to payment methods other than Google Play's billing system, with an exception for "tax exempt donations" to validated tax-exempt organizations (e.g., 501(c)(3) in the US). AnkiDroid's fiscal host, Open Source Collective, holds IRS 501(c)(6) tax-exempt status and provided the determination letter. However, Google rejected this documentation, stating the organization is "not tax-exempt" and that donations must be for a "validated 501(c)(3) charitable organization or local equivalent." Google has not clarified why 501(c)(6) status is insufficient, despite its policy referencing "tax exempt donations" (organization status) rather than tax-deductible donations (donor benefit). AnkiDroid will remove the donation link from its Play Store build under protest, as it is the sole funding source for the volunteer-maintained app.

Commenters highlight a critical distinction: 501(c)(6) organizations are tax-exempt (pay no income tax) but donations to them are not tax-deductible for donors, whereas Google's policy appears to require 501(c)(3) status where donations are tax-deductible. Several note Google's pattern of removing apps over payment policy disputes, citing WireGuard's 2019 ejection. Critics argue this demonstrates the danger of monopolistic app store control and advocate for alternative distribution (F-Droid, PWAs, sideloading). Some reference the Epic v. Google ruling and suggest alternative billing programs, though these involve service fees. Others express frustration that a free, donation-dependent educational app is being forced to choose between funding and distribution, calling for regulatory intervention in digital infrastructure.

3. I trained a small transformer in 1.5hrs and it beats many LLMs

HN discussion (535 points, 146 comments)

The author describes training a small autoregressive transformer from scratch on a single RTX 5090 in 1.5 hours ($0.67) that achieves 44% on ARC-AGI-1 and 7% on ARC-2, matching or exceeding many LLMs and specialized architectures like TRM and HRM. The method uses test-time training on evaluation puzzle inputs (with labels hidden), converting each input-output pair into token sequences augmented with color and dihedral permutations. Key architectural choices include 3D RoPE positional embeddings, per-task learned embeddings, SwiGLU activations, RMSNorm, and the NorMuon optimizer. The loss is computed only on output tokens (supervised), and the two most common augmented predictions are submitted. Ablations show 3D RoPE and per-task embeddings are critical (dropping to ~24% without them), while removing input-token training slightly improves scores (40%→44%). Adding filtered ARC-2 tasks boosts data diversity without leakage. The author argues ARC is an ideal sample-efficiency benchmark—few samples, high-dimensional, metalearning structure, human-solvable, and accessible—and that transformer scaling alone can reach 65% without recursion, massive synthetic data, or huge compute. The article also rebuts criticisms about "training on test," data leakage, benchmark policy, and compares against LLM-based approaches that rely on offline pretraining and synthetic data.

Commenters debate whether the approach constitutes "training on test" since it trains on evaluation puzzle inputs (though not labels), with some viewing this as a gray area akin to studying past exams, while the author defends it as standard transductive metalearning. Several note the benchmark specialization concern: the model is tailored to ARC puzzles rather than general reasoning, and question generalization beyond the benchmark. The author clarifies this is not an LLM but a small transformer targeting sample efficiency, and that frontier models might be beatable by training from scratch. Agreement emerges that LLM progress on ARC is driven by post-training on synthetic data rather than general reasoning, with private benchmarks needed for fair comparison. Other discussions touch on the author's personal story (medical emergency), the definition of "general reasoning," and technical questions about ARC-AGI-3 performance and the feasibility of banning offline training entirely.

4. Ask HN: Who is hiring? (September 2026)

HN discussion (172 points, 187 comments)

This is the September 2026 edition of Hacker News' monthly "Who is hiring?" thread, a community-driven job board where companies post open positions directly. The post establishes guidelines: only hiring companies (not recruiters) may post, one post per company, must include location and remote/onsite status, and companies should explain what they do if not well-known. The thread also links to third-party aggregation tools (hn-who-is-hiring, hnjobs, etc.) and references the companion "Who wants to be hired?" thread for job seekers.

The comments contain 20+ job postings from companies across North America, Europe, and globally remote roles. Notable postings include Modash (creator marketing platform, remote Europe, €75-110k), Quill (YC-backed analytics SDK, remote US, $150-210k), OysterHR (global employment platform, remote EMEA), Shepherd (AI-native commercial insurance, Series B, $42M raised, onsite SF/NYC), and Neuralwatt (AI datacenter energy optimization, remote Seattle/Denver, $180-220k). Several companies emphasize high ownership, small teams, and technical depth (Rust, distributed systems, ML). Roles span senior engineering, founding research, DevOps, and sales/operations, with compensation ranging from ~€60-70k/month in Amsterdam to $200-250k+equity in SF.

5. Introducing Ad Blocker for Firefox on iOS

HN discussion (259 points, 89 comments)

Mozilla has introduced a built-in Ad Blocker for Firefox on iOS that uses Apple's WebKit Content Blocker technology and the EasyList filter list to block third-party ads and ad-related trackers before they load. The feature is optional and off by default, accessible via Settings > Browsing > Ad Blocker. It does not block first-party ads served directly by websites, search engine result ads, or Firefox's own sponsored new-tab content. The ad blocker works alongside existing protections like Enhanced Tracking Protection. Mozilla notes that iOS lacks the extension ecosystem available on Desktop and Android, necessitating a built-in solution. The rollout is progressive and experimental, requiring "remote improvements" (telemetry) to be enabled for the feature to appear.

HN commenters express frustration that the feature remains unavailable to many users days after announcement due to a progressive rollout, and criticize the requirement to enable telemetry ("remote improvements") for ad blocking to function. Multiple users confirm it does not block YouTube ads or Reddit inline ads, and question why search engine ads are excluded—speculating this may relate to Mozilla's Google search deal. Comparisons favor Orion browser, which supports uBlock Origin via extensions on iOS, and Safari with uBlock Origin Lite. Technical questions arise about whether Firefox now supports the Content Blocker API for third-party tools like AdGuard, how this differs from the years-old Firefox Focus content blocker, and why Firefox doesn't use its own rendering engine on iOS (at least in the EU).

6. The ChatGPT/Codex app bundles a full copy of LibreOffice

HN discussion (183 points, 96 comments)

The author discovered that the OpenAI Codex desktop application (since rebranded as ChatGPT) stores approximately 1.7GB of data in `~/.cache/codex-runtimes/codex-primary-runtime/`. This cache includes full installations of Python and Node.js, along with native binaries for Poppler, git, and the complete LibreOffice office suite. The application also contains plugin skills that instruct Codex on how to locate and utilize these bundled binaries, suggesting they are used for document processing tasks such as reading, converting, and manipulating Microsoft Office files.

Commenters debated the necessity and implications of bundling LibreOffice. Several users speculated it enables headless conversion between document formats (docx, xlsx, pdf) for the new "Work" product, while others criticized the 1.7GB footprint as bloat, suggesting lighter WASM-based alternatives or on-demand downloading of hashed components. Some defended the approach, noting LibreOffice's reliability with legacy formats like old .xls files. Licensing concerns were raised regarding potential MPL 2.0 violations due to missing attributions. Comparisons were drawn to Claude Code's 10GB VM installation, with multiple users expressing frustration over the growing size and opacity of frontier AI desktop applications.

7. The creator of Jujutsu has joined ERSC

HN discussion (156 points, 120 comments)

East River Source Control (ERSC) has appointed Martin von Zweigbergk, creator of the Jujutsu (JJ) version control system, as chief technology officer. Von Zweigbergk began Jujutsu as a side project in 2019 before working on it full-time at Google; the project now has over 30,000 GitHub stars and is licensed under Apache 2.0. His prior experience includes work on Fig (a Mercurial client for Google's Piper monorepo) and contributions to Git. ERSC, launched in 2025 and backed by Amplify Partners, is building next-generation version control platforms for humans and machines, with ERSC Storage entering private beta in September 2026. Von Zweigbergk will continue as a core maintainer of Jujutsu. He stated that while Jujutsu improves the client-side experience, Git's server-side storage layer has scaling limitations that require a new storage model better supported by a company than an open-source project.

Commenters expressed skepticism about ERSC's value proposition versus existing solutions like GitHub, SourceHut, and Codeberg, questioning what surplus value ERSC provides given Jujutsu already works with Git. Several users clarified that ERSC appears to be developing an alternative backend storage layer for Jujutsu to replace Git's scaling limitations. Jujutsu received praise for its non-destructive workflow, undo capabilities, and improved UX over Git. Concerns were raised about the future of Jujutsu as an open-source project, while multiple commenters criticized ERSC's website for poor scroll performance, an animated background causing accessibility issues, and broken browser navigation gestures. Some noted the announcement appeared on LinkedIn weeks earlier, and one commenter questioned the "human ↔ machine collaboration" framing as "creepy."

8. Ambient CSS v3 – Blender meets CSS

HN discussion (173 points, 64 comments)

Unable to fetch article: No content extracted (possible paywall or JS-heavy site)

The discussion reveals a split between aesthetic appreciation and severe usability criticism. Many commenters enjoy the tactile, skeuomorphic aesthetic—comparing it to VST plugins, Teenage Engineering hardware, and Web 2.0-era interfaces—and welcome a departure from flat design. However, the implementation draws heavy fire: the demo site has broken scrolling, inconsistent light sources, laggy performance, and visual bugs (textures rendering as simple gradients, colors shifting unexpectedly). More critically, interactive components fail basic accessibility and UX standards—radio buttons and knobs are built with non-semantic `

`s and JavaScript instead of native inputs, breaking keyboard navigation and screen readers; knob rotation behaviors are inconsistent; and pressed/selected states are nearly indistinguishable. Several note the cycle of skeuomorphism returning as fashion, while others argue virtual knobs on flat screens are fundamentally flawed UX. The consensus is that while visually striking as a tech demo, the library is currently unsuitable for production use due to accessibility gaps, inconsistent interactions, and technical roughness.


Generated with hn-summaries