Create Missing RSS Feeds with LLMs

Comments (1)

PaulHoule · 66d ago

The general story about the LLM-scraper problem is that (1) "companies like OpenAI run badly implemented web crawlers to get training data" but there is (2) with LLMs scrapers could do content understanding (inference) that would make them more useful and I think the even more impactful (3) LLMs will empower people to write scrapers that would never have written them before.

I kinda laugh at (3) because it's been a running gag for me that management vastly overestimates the effort to write scrapers and crawlers because they've been burned with vastly underestimating the effort to develop what look like simple UI applications.

They usually think "this will be a hassle to maintain" but it usually isn't because: (a) the target web sites usually never change in a significant way because UI development is such a hassle and (b) the target web sites usually never change in a significant way because Google will punish them if they do [1]

It is like 10 minutes to write a scraper if you do it all the time and have an API like beautifulsoup on your fingertips, probably 20 minutes to vibe code it if you don't.

I am still using the same HTML scraper to process image galleries today that I used to process Flickr galleries back in the 00's, for a while the pattern was "fight with the OAuth to log into an API for 45 minutes" or "spend weeks figuring out how to parse MediaWiki markup" and then "get the old scraper working in less than 15 minutes". Frequently the scraper works perfectly out of the box, sometimes it works 80% out of the box, always it works 100% by adding a handful of rules.

I work on a product that has a React-based site and it seems the "state of the art" in scraping a URL [2] like

   https://example.com/item/8788481

is to download the HTML and then all the Javascript and CSS and other stuff with no cache (for every freaking page) and run the Javascript and have something scrape the content out of the DOM whereas they could just go to

   https://example.com/api/item/8788481

and get the data they want in a JSON format which could be processed like item["metadata"]["title"] or just stuffed into a JSONB column and queries any way you like. Login is not "fight with OAuth" but something like "POST username and password to https://example.com/api/login with a client that has a cookie jar" I don't really think "most people are stupid" that often but I think it all the time when web scraping is involved.

[1] they even have a patent for it! people who run online ad campaigns A/B test anything, but the last thing Google wants is for an SEO to be able to settle questions like "will my site rank higher if I put a certain phrase in a <b>?"

[2] ... as in, we see people doing it in our logs

LIGO detects most massive black hole merger to date (caltech.edu)

PHP License Update (wiki.php.net)

Kiro: A new agentic IDE (kiro.dev)

NeuralOS: An operating system powered by neural networks (neural-os.com)

Cognition (Devin AI) to Acquire Windsurf (cognition.ai)

Anthropic, Google, OpenAI and XAI Granted Up to $200M from Defense Department (cnbc.com)

Replicube: 3D shader puzzle game, online demo (replicube.xyz)

Building Modular Rails Applications: A Deep Dive into Rails Engines (panasiti.me)

Embedding user-defined indexes in Apache Parquet (datafusion.apache.org)

Cidco MailStation as a Z80 Development Platform (2019) (jcs.org)

Strategies for Fast Lexers (xnacly.me)

Context Rot: How increasing input tokens impacts LLM performance (research.trychroma.com)

Japanese grandparents create life-size Totoro with bus stop for grandkids (2020) (mymodernmet.com)

Lightning Detector Circuits (techlib.com)

Show HN: The HTML Maze - Escape an eerie labyrinth built with HTML pages (htmlmaze.com)

Meticulous (YC S21) is hiring in UK to redefine software dev (tinyurl.com)

SQLite async connection pool for high-performance (github.com)

Data brokers are selling flight information to CBP and ICE (eff.org)

Show HN: Bedrock – An 8-bit computing system for running programs anywhere (benbridle.com)

Two guys hated using Comcast, so they built their own fiber ISP (arstechnica.com)

The Corset X-Rays of Dr Ludovic O'Followell (1908) (publicdomainreview.org)

It took 45 years, but spreadsheet legend Mitch Kapor finally got his MIT degree (bostonglobe.com)

East Asian aerosol cleanup has likely contributed to global warming (nature.com)

Impacts of adding PV solar system to internal combustion engine vehicles (jstor.org)

Tandy Corporation, Part 3 Becoming IBM Compatible (abortretry.fail)

Show HN: Refine – A Local Alternative to Grammarly (refine.sh)

Lossless Float Image Compression (aras-p.info)

A Century of Quantum Mechanics (home.cern)

Six Game Devs Speak to Computer Games Mag (1984) (computeradsfromthepast.substack.com)

Why random selection is necessary to create stable meritocratic institutions (assemblingamerica.substack.com)

You Are in a Box (jyn.dev)

Happy 20th Birthday, Django (djangoproject.com)

A bionic knee integrated into tissue can restore natural movement (news.mit.edu)

Panasonic opens country's largest EV battery plant in De Soto, Kansas (kctv5.com)

AI slows down open source developers. Peter Naur can teach us why (johnwhiles.com)

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs (arxiv.org)

Self-imposed ban – a lightweight bash script to block commands (github.com)

GM, LG to upgrade Tennessee plant to make low-cost EV batteries (cnbc.com)

Lasagna Battery Cell (amazingribs.com)

Apple's Browser Engine Ban Persists, Even Under the DMA (open-web-advocacy.org)

Is there a cost to try catch blocks? (brandewinder.com)

Oakland cops gave ICE license plate data; SFPD also illegally shared with feds (sfstandard.com)

『 0x61 』- Panasonic and OpenBSD = <3 (x61.sh)

Kira Vale, $500 and 600 prompts, AI generated short movie [video] (youtube.com)

Concurrent Programming with Harmony (harmony.cs.cornell.edu)

Anthropic signs a $200M deal with the Department of Defense (anthropic.com)

Show HN: Ten years of running every day, visualized (nodaysoff.run)

OpenCut: The open-source CapCut alternative (github.com)

Binding Application in Idris (andrevidela.com)

The underground cathedral protecting Tokyo from floods (2018) (bbc.com)

Create Missing RSS Feeds with LLMs

Comments (1)