Files

11 KiB
Raw Permalink Blame History

name, description, version, tags
name description version tags
1688-cross-border-sourcing Scrape 1688.com for cross-border e-commerce products using Firecrawl API, rank by suitability, and enrich with Temu/Amazon pricing estimates. 1.0
1688
sourcing
cross-border
firecrawl
scraping

1688 Cross-Border E-Commerce Product Sourcing

Scrape 1688.com for products suitable for cross-border e-commerce (Temu/Amazon/AliExpress), then rank and analyze them.

⚠️ Critical: 1688 Anti-Scraping Reality

1688.com (owned by Alibaba) has extremely aggressive anti-bot measures. Most scraping methods FAIL:

Method Category Pages Detail Pages (offer/XXX) Reason
Direct browser_navigate Sandbox can't render 1688 JS
Jina Reader (r.jina.ai) Blocked by Alibaba Cloud WAF
Scrapling StealthyFetcher IP blocked by Alibaba Cloud
CamouFox (headless) Server IP blocked (datacenter)
Free proxy + CamouFox ⚠️ ⚠️ Page loads but login popup blocks
Google/Bing cache JS-only challenge page returned
AliExpress reference URLs JS-only SPA, 2KB bootloader shell
Firecrawl API Category/search listing pages work; detail pages TIMEOUT (2026-07-06 verified, 120s+ no response)
Desktop CDP Chrome ONLY method for individual detail pages — Chinese residential IP bypasses WAF

For category/search listing pages: Use Firecrawl API. For individual detail pages (detail.1688.com/offer/...): Desktop CDP is the ONLY reliable method. Do NOT attempt Firecrawl, Jina, or direct HTTP — they all fail from non-Chinese IPs.

Why CamouFox Alone Cannot Access 1688

1688's anti-bot runs on two independent layers (see skill:cloakbrowser "反爬两层防御模型"):

  1. Browser fingerprint layer — CamouFox bypasses this naturally
  2. IP reputation layer — Alibaba Cloud WAF rejects all known datacenter IPs outright. The response is a redirect to bixi.alicdn.com/punish/... with "Access denied" / "unusual traffic" message

Even with geoip=True and humanize=True, CamouFox on a server IP will be blocked at layer 2. The only way to browse 1688 interactively is CamouFox + Chinese residential proxy:

with Camoufox(headless=True, geoip=True, humanize=True,
              proxy={"server": "http://residential-proxy:port", "username": "user", "password": "pass"}) as browser:
    page = browser.new_page()
    page.goto("https://www.1688.com", timeout=30000)

For scraping (no interactive browsing needed), Firecrawl remains the best option — it routes through its own proxy infrastructure.

Step-by-Step Workflow

1. Scrape 1688 Category Pages via Firecrawl

Use Firecrawl API with waitFor=5000ms to extract product data. Focus on category listing pages, NOT search pages (search forces login).

import requests

FIRECRAWL_API_KEY = "fc-5536880adfa44fb78a3bd9da8dd52b53"

# Recommended categories for cross-border sourcing:
CATEGORIES = {
    "日用百货": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境日用百货",
    "宠物用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境宠物用品",
    "厨房用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境厨房小工具",
    "美妆工具": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境美妆工具",
    "家居收纳": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境家居收纳",
    "运动户外": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境运动户外",
    "数码配件": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境数码配件",
}

resp = requests.post("https://api.firecrawl.dev/v1/scrape", json={
    "url": category_url,
    "formats": ["markdown"],
    "waitFor": 5000,
}, headers={
    "Authorization": f"Bearer {FIRECRAWL_API_KEY}",
    "Content-Type": "application/json",
})
markdown_content = resp.json()["data"]["markdown"]

2. Parse Product Data from Markdown

Use regex to extract structured product data from Firecrawl's markdown output:

import re

# Price pattern
price_pattern = r'[¥¥]\s*(\d+\.?\d*)'
# Sales pattern: 成交X万+件 or 成交X+件
sales_pattern = r'成交(\d+\.?\d*)(万?)\+?件'
# Cross-border tags to detect
cross_border_tags = ['跨境', 'temu', '亚马逊', 'amazon', '专供', '爆款']

Products appear in markdown as sections with title, price, and sales info.

3. Score Products for Cross-Border Suitability

Composite scoring (0-100):

Factor Weight Criteria
Cross-border tag 30% Has "跨境/Temu/亚马逊" in title = full points
Sales volume 25% 10万+ = 25, 1万+ = 20, 1000+ = 15, <1000 = 10
Price point 25% ¥2-15 optimal range (cheap enough for margin)
Shipping friendliness 20% Light (<200g) + small package = full points

4. Enrich with Cross-Border Pricing Estimates

For each selected product, estimate:

  • TEMU price: 1688 price × 2.5-4x (USD), capped at typical TEMU range
  • Amazon price: 1688 price × 4-6x (USD), capped at typical Amazon range
  • Margin: (Selling price - 1688 cost - shipping) / Selling price
  • Typical shipping cost: Small packet <500g ≈ ¥15-25 to US/EU

5. Best Product Categories for Cross-Border

Based on analysis, these categories consistently perform well:

  1. 宠物用品 (Pet supplies) — High demand, emotional buying, good margins
  2. 厨房小工具 (Kitchen gadgets) — Lightweight, impulse buy, proven sellers
  3. 美妆工具 (Beauty tools) — Tech premium, high Amazon prices
  4. 家居收纳 (Home organizers) — Low cost, high utility, evergreen
  5. 日用百货 (Daily goods) — Volume play, staple items

Desktop CDP Browser — PRIMARY method for individual detail pages

⚠️ SG5 Server-Side Scraping: ALL Methods Blocked (2026-07 Verified)

SG5 server IP (43.160.244.125) is blocked by Alibaba Cloud WAF at the IP layer. These ALL fail:

Method Result Detail
Direct HTTP (curl/urllib) Redirect to bixi.alicdn.com/punish/...
Jina Reader (r.jina.ai) "Access denied" page
Firecrawl API ⚠️ Works ~50% but times out on JS-heavy detail pages
Google/Bing cache Returns search pages, not cached product data
AliExpress pages SPA only (2KB JS bootloader), no server-rendered data
AliExpress API endpoints x5sec anti-bot challenge on all endpoints

The ONLY reliable path: Desktop CDP Chrome (Chinese residential IP) → Bridge → scrape detail pages.

Primary: Desktop CDP Chrome Scrape (Chinese IP)

The user's Desktop runs Chrome on a Chinese residential IP, bypassing Alibaba Cloud's datacenter IP block entirely.

Automated scraping script: ~/.hermes/scripts/1688_scrape_cdp.py

  • Polls Bridge /health for Desktop slot with cdp_browser_running
  • Navigates to each 1688 offer URL via /cdp/navigate
  • Extracts title, images, price, specs via /cdp/evaluate
  • Writes enriched data to MongoDB premiumproducts.products
  • Deploy via cron: cronjob create no_agent=true script=1688_scrape_cdp.py schedule="every 5m"

Product Detail Enrichment Pipeline (Submit → Full Data)

For products submitted via Desktop Submit:

Step What Tool
1. Scrape Extract raw data from 1688 page CDP browser_console
2. Structure LLM extracts title/price/images/specs Qwen 35B (:40006) / qwen3-vl
3. Translate zh → en, zh → ru Qwen 35B / API
4. Images Remove watermark, resize 800×800 SDXL/FLUX + AtomK toolbox
5. Classify Auto-categorize product LLM
6. Write Update MongoDB document Server endpoint

Reference Files

  • references/cdp-detail-extraction.md — Extract structured product data from 1688 detail pages via Desktop CDP (body.innerText, image URL extraction, expression filter workarounds)
  • references/image-download-conversion.md — Download alicdn product images and convert WebP→JPG in Cloud Bridge sessions (cron no_agent pattern)
  • references/mongodb-product-update.md — Update product pool data via direct MongoDB connection when REST API is read-only (credentials, field schema, pitfalls)

Scripts

  • scripts/1688_scrape_cdp.py — Automated 1688 detail page scraper: health check → find Desktop slot → navigate → extract title/price/images/specs/weight → write MongoDB. Designed for cron no_agent=true polling: silently exits when no Desktop connected, runs when Desktop comes online.

Pitfalls

  • Do NOT use search pages (s.1688.com with keyword search forces login popup)
  • Category pages work better — products visible without login
  • Product detail links (offer/XXXX format) use encrypted redirects, hard to extract direct URLs
  • CamouFox on server IP = still blocked — 1688 blocks at IP layer (datacenter detection), not browser fingerprint. CamouFox alone is insufficient; must pair with residential proxy or use Firecrawl
  • Free proxies are unstable — may work one moment, blocked the next
  • Firecrawl waitFor=5000 is critical — without it, 1688's lazy-loaded content won't render
  • 🔴 Firecrawl detail page TIMEOUT (2026-07-06 verified): Firecrawl API times out (120s+) on detail.1688.com/offer/XXX URLs. It works for category/search listing pages only. For individual detail pages, use Desktop CDP (Chinese IP).
  • 🔴 AliExpress reference URLs are JS-only SPAs: aliexpress.com/item/XXX.html returns a 2KB JS bootloader with no product data. The internal API (aeglodetailweb/api/header) also blocks non-browser requests with x5sec protection. AliExpress cannot be used as a fallback data source from server-side code.
  • 🔴 Desktop must be connected for CDP scraping: When Desktop is not running (0 slots on Bridge /health), CDP scraping is unavailable. Use a cron polling script (no_agent=true) that silently exits when no Desktop is connected, and runs the scraper when it comes online. See scripts/1688_scrape_cdp.py for the canonical pattern.
  • Firecrawl API key can expire — returns {"success": false, "error": "Unauthorized: Invalid token"}. Fall back to Desktop CDP browser (Chinese IP) or renew the key.
  • Server-side CDP browser gets blockedbrowser_navigate to 1688 from SG3/SG5 returns bixi.alicdn.com/punish/... (Access Denied). Only Desktop Chrome on Chinese residential IP works for CDP-based scraping.
  • MOQ info not reliably available from category pages; need detail page access for that
  • Prices shown may be tier prices — the lowest price often requires high MOQ
  • Sales figures on 1688 can be inflated; treat as directional, not exact

Output Format

Save results as:

  • /tmp/1688_cross_border_top10.json — Full enriched data
  • /tmp/1688_cross_border_top10.csv — Spreadsheet-friendly format

CSV columns: 排名, 品类, 产品名称, 1688价格(元), 1688销量, 预估重量, TEMU参考价(), Amazon参考价(), 预估毛利率, 跨境标签, 推荐指数

  • web-scraping-toolkit — General scraping tools reference (Firecrawl is one of 5 tools documented there)
  • atomk-api-export — Alternative product sourcing from AtomK system
  • temu-pipeline — Downstream TEMU order processing