11 KiB
name, description, version, tags
| name | description | version | tags | |||||
|---|---|---|---|---|---|---|---|---|
| 1688-cross-border-sourcing | Scrape 1688.com for cross-border e-commerce products using Firecrawl API, rank by suitability, and enrich with Temu/Amazon pricing estimates. | 1.0 |
|
1688 Cross-Border E-Commerce Product Sourcing
Scrape 1688.com for products suitable for cross-border e-commerce (Temu/Amazon/AliExpress), then rank and analyze them.
⚠️ Critical: 1688 Anti-Scraping Reality
1688.com (owned by Alibaba) has extremely aggressive anti-bot measures. Most scraping methods FAIL:
| Method | Category Pages | Detail Pages (offer/XXX) | Reason |
|---|---|---|---|
| Direct browser_navigate | ❌ | ❌ | Sandbox can't render 1688 JS |
| Jina Reader (r.jina.ai) | ❌ | ❌ | Blocked by Alibaba Cloud WAF |
| Scrapling StealthyFetcher | ❌ | ❌ | IP blocked by Alibaba Cloud |
| CamouFox (headless) | ❌ | ❌ | Server IP blocked (datacenter) |
| Free proxy + CamouFox | ⚠️ | ⚠️ | Page loads but login popup blocks |
| Google/Bing cache | ❌ | ❌ | JS-only challenge page returned |
| AliExpress reference URLs | — | ❌ | JS-only SPA, 2KB bootloader shell |
| Firecrawl API | ✅ | ❌ | Category/search listing pages work; detail pages TIMEOUT (2026-07-06 verified, 120s+ no response) |
| Desktop CDP Chrome | — | ✅ | ONLY method for individual detail pages — Chinese residential IP bypasses WAF |
For category/search listing pages: Use Firecrawl API. For individual detail pages (detail.1688.com/offer/...): Desktop CDP is the ONLY reliable method. Do NOT attempt Firecrawl, Jina, or direct HTTP — they all fail from non-Chinese IPs.
Why CamouFox Alone Cannot Access 1688
1688's anti-bot runs on two independent layers (see skill:cloakbrowser "反爬两层防御模型"):
- Browser fingerprint layer — CamouFox bypasses this naturally
- IP reputation layer — Alibaba Cloud WAF rejects all known datacenter IPs outright. The response is a redirect to
bixi.alicdn.com/punish/...with "Access denied" / "unusual traffic" message
Even with geoip=True and humanize=True, CamouFox on a server IP will be blocked at layer 2. The only way to browse 1688 interactively is CamouFox + Chinese residential proxy:
with Camoufox(headless=True, geoip=True, humanize=True,
proxy={"server": "http://residential-proxy:port", "username": "user", "password": "pass"}) as browser:
page = browser.new_page()
page.goto("https://www.1688.com", timeout=30000)
For scraping (no interactive browsing needed), Firecrawl remains the best option — it routes through its own proxy infrastructure.
Step-by-Step Workflow
1. Scrape 1688 Category Pages via Firecrawl
Use Firecrawl API with waitFor=5000ms to extract product data. Focus on category listing pages, NOT search pages (search forces login).
import requests
FIRECRAWL_API_KEY = "fc-5536880adfa44fb78a3bd9da8dd52b53"
# Recommended categories for cross-border sourcing:
CATEGORIES = {
"日用百货": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境日用百货",
"宠物用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境宠物用品",
"厨房用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境厨房小工具",
"美妆工具": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境美妆工具",
"家居收纳": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境家居收纳",
"运动户外": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境运动户外",
"数码配件": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境数码配件",
}
resp = requests.post("https://api.firecrawl.dev/v1/scrape", json={
"url": category_url,
"formats": ["markdown"],
"waitFor": 5000,
}, headers={
"Authorization": f"Bearer {FIRECRAWL_API_KEY}",
"Content-Type": "application/json",
})
markdown_content = resp.json()["data"]["markdown"]
2. Parse Product Data from Markdown
Use regex to extract structured product data from Firecrawl's markdown output:
import re
# Price pattern
price_pattern = r'[¥¥]\s*(\d+\.?\d*)'
# Sales pattern: 成交X万+件 or 成交X+件
sales_pattern = r'成交(\d+\.?\d*)(万?)\+?件'
# Cross-border tags to detect
cross_border_tags = ['跨境', 'temu', '亚马逊', 'amazon', '专供', '爆款']
Products appear in markdown as sections with title, price, and sales info.
3. Score Products for Cross-Border Suitability
Composite scoring (0-100):
| Factor | Weight | Criteria |
|---|---|---|
| Cross-border tag | 30% | Has "跨境/Temu/亚马逊" in title = full points |
| Sales volume | 25% | 10万+ = 25, 1万+ = 20, 1000+ = 15, <1000 = 10 |
| Price point | 25% | ¥2-15 optimal range (cheap enough for margin) |
| Shipping friendliness | 20% | Light (<200g) + small package = full points |
4. Enrich with Cross-Border Pricing Estimates
For each selected product, estimate:
- TEMU price: 1688 price × 2.5-4x (USD), capped at typical TEMU range
- Amazon price: 1688 price × 4-6x (USD), capped at typical Amazon range
- Margin: (Selling price - 1688 cost - shipping) / Selling price
- Typical shipping cost: Small packet <500g ≈ ¥15-25 to US/EU
5. Best Product Categories for Cross-Border
Based on analysis, these categories consistently perform well:
- 宠物用品 (Pet supplies) — High demand, emotional buying, good margins
- 厨房小工具 (Kitchen gadgets) — Lightweight, impulse buy, proven sellers
- 美妆工具 (Beauty tools) — Tech premium, high Amazon prices
- 家居收纳 (Home organizers) — Low cost, high utility, evergreen
- 日用百货 (Daily goods) — Volume play, staple items
Desktop CDP Browser — PRIMARY method for individual detail pages
⚠️ SG5 Server-Side Scraping: ALL Methods Blocked (2026-07 Verified)
SG5 server IP (43.160.244.125) is blocked by Alibaba Cloud WAF at the IP layer. These ALL fail:
| Method | Result | Detail |
|---|---|---|
| Direct HTTP (curl/urllib) | ❌ | Redirect to bixi.alicdn.com/punish/... |
| Jina Reader (r.jina.ai) | ❌ | "Access denied" page |
| Firecrawl API | ⚠️ | Works ~50% but times out on JS-heavy detail pages |
| Google/Bing cache | ❌ | Returns search pages, not cached product data |
| AliExpress pages | ❌ | SPA only (2KB JS bootloader), no server-rendered data |
| AliExpress API endpoints | ❌ | x5sec anti-bot challenge on all endpoints |
The ONLY reliable path: Desktop CDP Chrome (Chinese residential IP) → Bridge → scrape detail pages.
Primary: Desktop CDP Chrome Scrape (Chinese IP)
The user's Desktop runs Chrome on a Chinese residential IP, bypassing Alibaba Cloud's datacenter IP block entirely.
Automated scraping script: ~/.hermes/scripts/1688_scrape_cdp.py
- Polls Bridge
/healthfor Desktop slot withcdp_browser_running - Navigates to each 1688 offer URL via
/cdp/navigate - Extracts title, images, price, specs via
/cdp/evaluate - Writes enriched data to MongoDB
premiumproducts.products - Deploy via cron:
cronjob create no_agent=true script=1688_scrape_cdp.py schedule="every 5m"
Product Detail Enrichment Pipeline (Submit → Full Data)
For products submitted via Desktop Submit:
| Step | What | Tool |
|---|---|---|
| 1. Scrape | Extract raw data from 1688 page | CDP browser_console |
| 2. Structure | LLM extracts title/price/images/specs | Qwen 35B (:40006) / qwen3-vl |
| 3. Translate | zh → en, zh → ru | Qwen 35B / API |
| 4. Images | Remove watermark, resize 800×800 | SDXL/FLUX + AtomK toolbox |
| 5. Classify | Auto-categorize product | LLM |
| 6. Write | Update MongoDB document | Server endpoint |
Reference Files
references/cdp-detail-extraction.md— Extract structured product data from 1688 detail pages via Desktop CDP (body.innerText, image URL extraction, expression filter workarounds)references/image-download-conversion.md— Download alicdn product images and convert WebP→JPG in Cloud Bridge sessions (cron no_agent pattern)references/mongodb-product-update.md— Update product pool data via direct MongoDB connection when REST API is read-only (credentials, field schema, pitfalls)
Scripts
scripts/1688_scrape_cdp.py— Automated 1688 detail page scraper: health check → find Desktop slot → navigate → extract title/price/images/specs/weight → write MongoDB. Designed for cronno_agent=truepolling: silently exits when no Desktop connected, runs when Desktop comes online.
Pitfalls
- Do NOT use search pages (s.1688.com with keyword search forces login popup)
- Category pages work better — products visible without login
- Product detail links (offer/XXXX format) use encrypted redirects, hard to extract direct URLs
- CamouFox on server IP = still blocked — 1688 blocks at IP layer (datacenter detection), not browser fingerprint. CamouFox alone is insufficient; must pair with residential proxy or use Firecrawl
- Free proxies are unstable — may work one moment, blocked the next
- Firecrawl waitFor=5000 is critical — without it, 1688's lazy-loaded content won't render
- 🔴 Firecrawl detail page TIMEOUT (2026-07-06 verified): Firecrawl API times out (120s+) on
detail.1688.com/offer/XXXURLs. It works for category/search listing pages only. For individual detail pages, use Desktop CDP (Chinese IP). - 🔴 AliExpress reference URLs are JS-only SPAs:
aliexpress.com/item/XXX.htmlreturns a 2KB JS bootloader with no product data. The internal API (aeglodetailweb/api/header) also blocks non-browser requests with x5sec protection. AliExpress cannot be used as a fallback data source from server-side code. - 🔴 Desktop must be connected for CDP scraping: When Desktop is not running (0 slots on Bridge /health), CDP scraping is unavailable. Use a cron polling script (
no_agent=true) that silently exits when no Desktop is connected, and runs the scraper when it comes online. Seescripts/1688_scrape_cdp.pyfor the canonical pattern. - Firecrawl API key can expire — returns
{"success": false, "error": "Unauthorized: Invalid token"}. Fall back to Desktop CDP browser (Chinese IP) or renew the key. - Server-side CDP browser gets blocked —
browser_navigateto 1688 from SG3/SG5 returnsbixi.alicdn.com/punish/...(Access Denied). Only Desktop Chrome on Chinese residential IP works for CDP-based scraping. - MOQ info not reliably available from category pages; need detail page access for that
- Prices shown may be tier prices — the lowest price often requires high MOQ
- Sales figures on 1688 can be inflated; treat as directional, not exact
Output Format
Save results as:
/tmp/1688_cross_border_top10.json— Full enriched data/tmp/1688_cross_border_top10.csv— Spreadsheet-friendly format
CSV columns: 排名, 品类, 产品名称, 1688价格(元), 1688销量, 预估重量, TEMU参考价(), Amazon参考价(), 预估毛利率, 跨境标签, 推荐指数
Related Skills
web-scraping-toolkit— General scraping tools reference (Firecrawl is one of 5 tools documented there)atomk-api-export— Alternative product sourcing from AtomK systemtemu-pipeline— Downstream TEMU order processing