The thing about buying small parts from AliExpress is that you don't lose the part. You lose the sheet that tells you what the part actually is.
The listing changes, the seller relists under a new ID, the whole page vanishes one day, and the only record of the thread pitch and the flow rate was in a table you scrolled past in April.
Daniel's been sitting on that problem for a while. He's got the home inventory system already, the one that records where things physically are, and what he wants to add is a snapshot function. Save the AliExpress link, pull the page, keep the part of it that reads like a technical sheet. No reviews. He doesn't even want the price.
That's the detail that makes this interesting.
He wants it targeted. Define what you extract up front, get only that back. And he's aware of the obvious failure mode there, which is that you define the divs, AliExpress redesigns the page, and your scraper is quietly extracting nothing.
So he's weighing two things. The traditional route, an Apify actor doing selector work, against an agentic extraction layer where a prompt tells the model what to pull. And then the part he flagged himself, which is that AliExpress is one of the most scraped sites on earth, so anti-bot is going to be a factor.
His ideal flow is worth keeping in view. He saves a URL against a catalog item, that triggers the runner, and if it works the PDF gets uploaded and appended to the item through his own internal API. So the extraction is one piece of a pipeline, not the whole thing.
Here's what I want to get to first, because it changes how you read his worry. He says he could define the divs he needs, but that risks breaking if AliExpress changes their front end.
That was the right fear. It's just late.
The AliExpress product page is client-side rendered now. The shell you get back, the raw HTML, is around seventy-seven kilobytes and contains no price and no SKU data. Everything is assembled in the browser after the fact.
And the piece that older tutorials lean on, the window dot run params object where the product data used to sit inline?
Empty object. It's still there in the page, it's just empty. Which is the nastiest version of this, because a parser that reads it doesn't throw an error. It returns nothing, cheerfully, and your job logs a success.
So it already broke. He was planning around a risk that has already materialized on the site he's targeting.
One of the stealth browser projects, running a fresh identity, got zero of twenty-four page loads on a flagged IP. The page either renders with Baxia's challenge in it or it doesn't render at all.
So the surface-level question, selectors or agent, is downstream of something bigger.
Considerably. And I think we should give him the real comparison anyway, because he needs it to make the call, and then we should show him why the comparison keeps collapsing. So if the old selector approach is dead on arrival, what are the two live options?
Option one. He mentioned Apify, which is the sensible place to look, because Apify's store has an AliExpress scraper that specifically returns what he wants.
The one to look at is by a builder called Goldmine, the logical underscore scrapers AliExpress scraper. It returns a specs array. Name and value pairs. That is the technical sheet, already parsed for you.
That's the exact field he described wanting.
There's an include description flag if he wants the description HTML as well, plus category path IDs, variants, shipping, store. And it handles proxies automatically, residential proxies in whatever country he's shipping to.
So the geography problem, the fact that the same product shows a different price and sometimes a different listing depending on where you're asking from, is handled inside the actor.
For price, yes, though he doesn't care about price. Pricing is pay per result. Three point nine nine tenths of a cent, roughly, per product on the free tier, dropping to two point nine nine on the paid tiers, so about two dollars ninety-nine per thousand products.
That's cheap enough that this is not a budget question. What's the reliability look like?
Four hundred and ninety-five total users, about eighty point nine percent of runs succeeded, four point nine eight out of five. And the FAQ on the actor page is unusually candid. It says AliExpress sometimes refuses product details to an IP address, the actor retries each product from several addresses, and a product it still cannot load keeps its listing fields with its detail fields empty.
Which is the same silent failure, just inside the black box.
Same shape. The run succeeds, the row exists, the specs field is empty.
The trade-off is that he doesn't own the maintenance. He inherits someone else's.
That's the honest version of the argument for this path. His fear about front-end changes is correct, but on this route that fear becomes the vendor's problem. The actor page says it's tested against live pages and fixed when the site changes. That's not a guarantee. It's a division of labor.
And the trigger pattern he described, run-on-save, maps cleanly. Apify fires webhooks on run events, and the only action available right now is a POST to whatever URL you specify.
So save the URL, POST to the actor, let the webhook POST back to his pipeline when the run finishes. That part is boring, which is a compliment.
Option two. The agentic layer, where the prompt does the extracting.
Firecrawl's JSON mode is the closest thing to what he described. You pass it a prompt, or a schema, or both, and it returns structured data from a single known URL. It's synchronous, it's one request, it doesn't need to navigate anything.
And the docs give the exact guidance someone building this needs. There's a line about adding location hints. The example they use is telling the model to pull flow rate in GPM from the specifications table.
That's the whole trick for a specs table. You're not just naming the field, you're naming where on the page it lives, because the model can actually read the page and find the table with that heading. Which is exactly what a CSS selector does, except it's resilient to the table moving.
There's a second piece of guidance that matters more than it sounds. Include null handling in the field descriptions, so the model doesn't guess missing values.
That's the one thing I'd tattoo on anyone building an extraction layer. A model asked for a flow rate when there isn't one will invent a plausible flow rate. You have to tell it, in the schema, that empty is an acceptable answer.
Does it handle the page attributes? The data IDs, the custom attributes?
No, and this is a v2 caveat worth knowing. HTML attributes aren't available in JSON extraction. The extraction runs on the markdown conversion of the page, so data attributes get stripped before the model ever sees them.
For a specs table that's fine. It's visible text, and visible text survives the markdown conversion.
Fine for this task, useless if he ever wanted to target something by attribute rather than by what it says.
And Firecrawl has an agent endpoint that goes further, autonomous navigation, no URL needed. Overkill here?
Their own docs say so. JSON mode on the scrape endpoint is cheaper and synchronous for one known URL. The agent is for when you don't know where the thing is.
Then there's the open-source option, Browser Use, where you describe the scraping task in plain language instead of writing selectors.
And this is where the picture gets complicated, because Browser Use's own CTO published a piece in September called The Bitter Lesson of Browser Agents. The argument is that pure LLM agents are slow, expensive, and unpredictable for high-frequency tasks.
That's a company arguing against the thing its own tool does.
It's a company following the evidence. Their answer is a product called Workflow Use. You record a workflow once, deterministically, and then it runs reliably a million times, and when the page changes underneath it, the model repairs the workflow rather than the workflow being rewritten by hand.
So record the reliable path, let the model patch it when it breaks.
That's the hybrid, and it directly answers his stated fear. He doesn't want to maintain selectors, and he doesn't want an agent burning tokens and time on every single run. Workflow Use is the middle.
The bitter lesson framing is honestly the most useful thing in this whole comparison. The industry's practitioners have already run the experiment he's proposing to run, and they concluded the same thing.
Both of those architectures assume you can actually reach the page.
That's where this gets interesting, and it's the part Daniel half-anticipated. Let's do the anti-bot layer properly, because I think it's the decisive thing and the rest is decoration.
AliExpress runs Alibaba's Baxia system. The public testing on it is fairly stark. Blocked requests come back with an HTTP two hundred status.
Two hundred. The success code.
Two hundred, with a bxpunish header and body markers. Things like rgv587 flag, x5secdata, fragments of the tmd string, FAIL SYS USER VALIDATE.
So if your dashboard counts two hundreds as successes, it's green while every row is empty.
Bright Data's line on it is the one I'd put on a wall. If a dashboard counts HTTP two-xx responses as successes, it reports a healthy pipeline while the rows are empty. That's not a scraping problem, that's a monitoring problem that hides a scraping problem.
How bad is the actual ceiling? If he builds this himself and just runs it, how far does he get?
There's public testing on the signed internal API, the mtop endpoint, the one that needs an MD5 signature built from a token, a timestamp, an app key, and the body. Even with a correct signature, from a single residential IP, one run got seven successes in sixty calls.
Seven.
Fifteen in sixty on a second run. The first block landed at request eight and request twelve respectively.
So you get maybe a dozen good pulls before Baxia decides you're a bot, and then it's over for that IP.
New sessions didn't clear it. Sixty-second waits didn't clear it. The blocked responses were exactly five hundred and forty-three bytes each, which is another tell. If you're logging response sizes, you can spot the block by the shape of it.
Right. So the choice of extraction library is competing for second place behind the choice of how you get the request to land at all.
By a wide margin. Bright Data's own conclusion is that you scrape the pages the site makes public and let your infrastructure handle IP addresses and fingerprints. Their three managed products all got through where self-run stealth browsers got zero out of twenty-four loads on a flagged IP.
And that's the honest answer to what he asked. He wanted a recommendation between two architectures. The recommendation is that he pick the unblocking layer first, then the extractor, and if he's picking a maintained actor, he's really picking a maintained unblocking layer with an extraction layer attached.
There's a smaller point in there worth pulling out, because he'd run into it. The framing of browserless.
He used that word.
The agentic tools that actually work against AliExpress all run a browser. Firecrawl, Browser Use, the Bright Data browser API. Browserless in the sense of no headless Chrome is not achievable against a client-side rendered page with Baxia in front of it. The browser is the thing doing the rendering.
The intelligence moved up a layer. The browser didn't go away.
And then there's the page economics, which I find funny. The mtop response for a product has twenty-seven top-level modules. The useful ones, the eight that carry data you'd want, are sixteen and a half kilobytes. About fifteen point nine percent of the payload.
What's the rest?
Shipping is the largest single module at thirty-four thousand bytes. Global data at twenty-three thousand. So you're paying bandwidth for shipping scaffolding you didn't ask for, on a page that renders to about three megabytes when it's fully built out.
So even a successful request is mostly freight.
And there's the surface-level scoring worth knowing. One anti-bot index rates AliExpress easy, one out of ten, at the homepage. But it flags that deep pages, profiles, listings, search, are usually harder, and product detail specifically is readily CAPTCHA-gated when the hits are rapid or the referer is missing.
Which is the whole design. Make the homepage look soft, make the deep data expensive.
So the architecture comparison, on its own terms, favors the hybrid. Targeted extraction with a maintained unblocking layer. But the anti-bot reality makes it less a preference and more the only thing that works.
There's the PDF piece too, which he kind of set aside and I don't think he should.
Neither Apify nor Firecrawl returns a PDF. They return JSON or markdown.
So the pipeline needs a render step. Markdown to PDF, somewhere in his stack, after extraction and before the append call.
Which is not hard, but it's a real component, and it's the one place where the tooling is his to own. The trigger maps onto Apify webhooks, or Firecrawl's agent-completed event if he goes that route. The extraction maps onto the actor or the JSON mode. The unblocking maps onto whoever's selling proxies. The render, the PDF, the append through his internal API, that's his.
And notably, nobody has built the thing he wants. No off-the-shelf AliExpress technical-sheet-to-PDF tool exists.
Right. The closest are general product scrapers that happen to include a specs array, and general LLM extraction endpoints. The snapshot-to-PDF-and-file-it pipeline is his to assemble.
There's one gap I want to flag honestly, because it changes what he should test first. Nobody's published whether the specs section specifically is served from that signed API or from the rendered DOM. The teardowns cover price, SKU, reviews. The specs block is not addressed.
Which matters. If specs come from the rendered DOM, he needs full rendering against Baxia every time. If specs come from the signed API, it's a lighter request, and his whole cost and rate-limit math shifts.
So the first thing he should do is open a product page with the network panel and see where the specs table actually comes from. That's an hour of work that determines the architecture more than the Apify versus agent question does.
And the specs table is the one part of the page that doesn't change much, which is exactly why it's worth keeping. The price and the promotion change hourly, the thread pitch doesn't. Which is also why caching a snapshot once per product is viable at all, and why his price-blindness is a feature rather than a limitation.
Every field he isn't scraping is a field he isn't fighting for. Price is the most volatile, most region-dependent, most aggressively protected thing on the page, and he's walked around it.
And that's the thing. The specs sheet is the one part of the page that doesn't change much, which is exactly why it's worth keeping.
You're both right about the sheet being the thing worth keeping. But you're both wrong about why.
You're scraping AliExpress because you don't trust AliExpress to still have the page. Fine. That's not a scraping problem. That's a "what happens when the catalogue goes out of print" problem, and I've watched that happen to a man whose whole system died with a single document.
Which man?
My uncle. Had a collection of small technical parts. Fittings, mostly, and a few hundred of them, boxed and labelled. And he had a system for it, which was index cards in a filing cabinet.
Cards, not a spreadsheet.
Cards. And the cards didn't say what the part was. They said the page number.
The page number of what?
Of the catalogue. One physical copy of an industrial supply catalogue, and it lived in his garage on a shelf above the workbench. The card gave you the part's drawer and the page where the specs were printed. That was the whole index. You found the card, you found the page, you had your thread pitch and your material and your pressure rating.
So the card was a pointer, not a record.
It worked beautifully. For eleven years. Then the catalogue went out of print, the supplier stopped sending new editions, and the copy in the garage was the only one he had. By the time he'd passed, that shelf was the only place those specs existed. The cards weren't wrong. They just pointed at nothing.
That's the actual argument for Daniel's PDF snapshot. Not convenience. Permanence.
The snapshot's the right instinct. Just not for the reason he gave. He thinks he's solving a scraper problem. He's solving a "the source will disappear" problem, and the source disappearing doesn't have to mean anti-bot. Sometimes it just means the catalogue stopped being printed.
The card's assumption was that the catalogue would always exist in a place he could reach. Same assumption as the URL.
Up to a point.
Can I ask what happened to the cards?
I have them.
You have them.
In a box. I've been meaning to digitise them for years. The problem is I can't find anyone who still owns that catalogue, so the page numbers don't mean anything to anyone but me, and I keep getting distracted by a dispute with my neighbour about a shared driveway.
The driveway.
He's been parking on my side since April. The cards are not going anywhere. The driveway is going somewhere.
The catalogue itself, did he buy it?
No. He won it.
He won a catalogue.
In a raffle. At a trade show, in a city I can't quite name. The prize was the catalogue, a leather-bound limited edition, and a lifetime supply of one specific size of O-ring, which he never used because it didn't fit any of his parts.
How many O-rings are we talking about?
I'd have to count. They're in the same box as the cards.
You have the O-rings.
I have the O-rings. And I've been reading up on AliExpress, because I'm going to list them. Which brings me to my question, which is about seller-side anti-bot measures, and I don't think either of you is prepared to answer it.
We are not.
We are not.
Then never mind.
The buyer-side scraping problem, and the seller-side listing problem, in the same house.
In the same box. That's the part nobody thinks about. The card is only as good as the thing it points at, and the thing it points at doesn't owe you anything. The catalogue goes out of print. The listing gets pulled. You keep the card and you keep the O-rings and you've got a very tidy record of something you can't look up anymore.
Which is why Daniel is right to want the sheet itself, and not a link to it.
Why the pointer-versus-record distinction is the thing I'm taking from this. The card and the URL are the same object. They're both promises that the underlying document will still be there.
An hour ago I'd have said the interesting question was Apify against the agent layer.
It's still worth doing the comparison. But the ceiling isn't the library. It's that Baxia gives you about a dozen good requests per IP before it stops answering.
Which means the choice is less "which tool" and more "whose unblocking layer do you trust, and what does it cost you per thousand." That's a procurement question wearing a technical costume.
It's a strange thing to realise about a personal project. He's building a small personal system, and the honest advice is to pay someone who's already solved the unblocking problem, rather than solve it himself for one user.
I'm not sure that's good news. The unblocking problem only gets solvable at scale, and scale means vendors, and vendors mean a handful of companies standing between you and the public web. His PDF snapshot is a way of not depending on AliExpress. But the thing that produces it depends on Bright Data or Apify or whoever, and those are fewer and fewer.
Which is the centralisation question, and it doesn't resolve cleanly. Every layer of the pipeline has somebody who owns it and somebody who rents access to it. His snapshot is a hedge, but it's a hedge purchased through a chain of other dependencies.
Still. The sheet is the thing worth keeping. He's right about that, and Hilbert's uncle is a decent argument for it.
The one concrete piece of advice I'd leave him with is the network panel question. Before he picks a stack, find out where the specs table comes from. If it's the signed API, the whole architecture is lighter and cheaper than we've been assuming, and the scraping comparison matters less than he thinks.
If it's the DOM, he's rendering through Baxia on every single product, and the unblocking layer is the project, not a component of it.
That's it.
We've kept you long enough on the driveway. Thanks to Hilbert Flumingtop, our producer, for being on the desk, and to Daniel for a prompt that turned out to be about archival more than it was about scraping. If you enjoyed this one, try episode eight, Building Your Own Whisper; episode two, Local STT For AMD GPU Owners; and episode twenty, Architectural AI. This has been My Weird Prompts. If you've enjoyed it, a review wherever you're listening helps people find the show. Send us your own prompt on Telegram at t dot me slash MWP listener bot. We'll be back soon.
Go check where your specs live.