Files

210 lines
11 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: 1688-cross-border-sourcing
description: Scrape 1688.com for cross-border e-commerce products using Firecrawl API, rank by suitability, and enrich with Temu/Amazon pricing estimates.
version: 1.0
tags: [1688, sourcing, cross-border, firecrawl, scraping]
---
# 1688 Cross-Border E-Commerce Product Sourcing
Scrape 1688.com for products suitable for cross-border e-commerce (Temu/Amazon/AliExpress), then rank and analyze them.
## ⚠️ Critical: 1688 Anti-Scraping Reality
1688.com (owned by Alibaba) has **extremely aggressive anti-bot measures**. Most scraping methods FAIL:
| Method | Category Pages | Detail Pages (offer/XXX) | Reason |
|--------|:---:|:---:|--------|
| Direct browser_navigate | ❌ | ❌ | Sandbox can't render 1688 JS |
| Jina Reader (r.jina.ai) | ❌ | ❌ | Blocked by Alibaba Cloud WAF |
| Scrapling StealthyFetcher | ❌ | ❌ | IP blocked by Alibaba Cloud |
| CamouFox (headless) | ❌ | ❌ | Server IP blocked (datacenter) |
| Free proxy + CamouFox | ⚠️ | ⚠️ | Page loads but login popup blocks |
| Google/Bing cache | ❌ | ❌ | JS-only challenge page returned |
| AliExpress reference URLs | — | ❌ | JS-only SPA, 2KB bootloader shell |
| **Firecrawl API** | ✅ | ❌ | Category/search listing pages work; **detail pages TIMEOUT** (2026-07-06 verified, 120s+ no response) |
| **Desktop CDP Chrome** | — | ✅ | **ONLY method for individual detail pages** — Chinese residential IP bypasses WAF |
**For category/search listing pages**: Use Firecrawl API.
**For individual detail pages (detail.1688.com/offer/...)**: Desktop CDP is the ONLY reliable method. Do NOT attempt Firecrawl, Jina, or direct HTTP — they all fail from non-Chinese IPs.
### Why CamouFox Alone Cannot Access 1688
1688's anti-bot runs on two independent layers (see skill:cloakbrowser "反爬两层防御模型"):
1. Browser fingerprint layer — CamouFox bypasses this naturally
2. IP reputation layer — Alibaba Cloud WAF rejects all known datacenter IPs outright. The response is a redirect to `bixi.alicdn.com/punish/...` with "Access denied" / "unusual traffic" message
Even with `geoip=True` and `humanize=True`, CamouFox on a server IP will be blocked at layer 2. The only way to browse 1688 interactively is CamouFox + Chinese residential proxy:
```python
with Camoufox(headless=True, geoip=True, humanize=True,
proxy={"server": "http://residential-proxy:port", "username": "user", "password": "pass"}) as browser:
page = browser.new_page()
page.goto("https://www.1688.com", timeout=30000)
```
For scraping (no interactive browsing needed), Firecrawl remains the best option — it routes through its own proxy infrastructure.
## Step-by-Step Workflow
### 1. Scrape 1688 Category Pages via Firecrawl
Use Firecrawl API with `waitFor=5000ms` to extract product data. Focus on **category listing pages**, NOT search pages (search forces login).
```python
import requests
FIRECRAWL_API_KEY = "fc-5536880adfa44fb78a3bd9da8dd52b53"
# Recommended categories for cross-border sourcing:
CATEGORIES = {
"日用百货": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境日用百货",
"宠物用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境宠物用品",
"厨房用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境厨房小工具",
"美妆工具": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境美妆工具",
"家居收纳": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境家居收纳",
"运动户外": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境运动户外",
"数码配件": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境数码配件",
}
resp = requests.post("https://api.firecrawl.dev/v1/scrape", json={
"url": category_url,
"formats": ["markdown"],
"waitFor": 5000,
}, headers={
"Authorization": f"Bearer {FIRECRAWL_API_KEY}",
"Content-Type": "application/json",
})
markdown_content = resp.json()["data"]["markdown"]
```
### 2. Parse Product Data from Markdown
Use regex to extract structured product data from Firecrawl's markdown output:
```python
import re
# Price pattern
price_pattern = r'[¥¥]\s*(\d+\.?\d*)'
# Sales pattern: 成交X万+件 or 成交X+件
sales_pattern = r'成交(\d+\.?\d*)(万?)\+?件'
# Cross-border tags to detect
cross_border_tags = ['跨境', 'temu', '亚马逊', 'amazon', '专供', '爆款']
```
Products appear in markdown as sections with title, price, and sales info.
### 3. Score Products for Cross-Border Suitability
Composite scoring (0-100):
| Factor | Weight | Criteria |
|--------|--------|----------|
| Cross-border tag | 30% | Has "跨境/Temu/亚马逊" in title = full points |
| Sales volume | 25% | 10万+ = 25, 1万+ = 20, 1000+ = 15, <1000 = 10 |
| Price point | 25% | ¥2-15 optimal range (cheap enough for margin) |
| Shipping friendliness | 20% | Light (<200g) + small package = full points |
### 4. Enrich with Cross-Border Pricing Estimates
For each selected product, estimate:
- **TEMU price**: 1688 price × 2.5-4x (USD), capped at typical TEMU range
- **Amazon price**: 1688 price × 4-6x (USD), capped at typical Amazon range
- **Margin**: (Selling price - 1688 cost - shipping) / Selling price
- **Typical shipping cost**: Small packet <500g ≈ ¥15-25 to US/EU
### 5. Best Product Categories for Cross-Border
Based on analysis, these categories consistently perform well:
1. **宠物用品** (Pet supplies) — High demand, emotional buying, good margins
2. **厨房小工具** (Kitchen gadgets) — Lightweight, impulse buy, proven sellers
3. **美妆工具** (Beauty tools) — Tech premium, high Amazon prices
4. **家居收纳** (Home organizers) — Low cost, high utility, evergreen
5. **日用百货** (Daily goods) — Volume play, staple items
## Desktop CDP Browser — PRIMARY method for individual detail pages
## ⚠️ SG5 Server-Side Scraping: ALL Methods Blocked (2026-07 Verified)
SG5 server IP (43.160.244.125) is blocked by Alibaba Cloud WAF at the IP layer. These ALL fail:
| Method | Result | Detail |
|--------|--------|--------|
| Direct HTTP (curl/urllib) | ❌ | Redirect to `bixi.alicdn.com/punish/...` |
| Jina Reader (r.jina.ai) | ❌ | "Access denied" page |
| Firecrawl API | ⚠️ | Works ~50% but times out on JS-heavy detail pages |
| Google/Bing cache | ❌ | Returns search pages, not cached product data |
| AliExpress pages | ❌ | SPA only (2KB JS bootloader), no server-rendered data |
| AliExpress API endpoints | ❌ | x5sec anti-bot challenge on all endpoints |
**The ONLY reliable path**: Desktop CDP Chrome (Chinese residential IP) → Bridge → scrape detail pages.
## Primary: Desktop CDP Chrome Scrape (Chinese IP)
The user's Desktop runs Chrome on a Chinese residential IP, bypassing Alibaba Cloud's datacenter IP block entirely.
**Automated scraping script**: `~/.hermes/scripts/1688_scrape_cdp.py`
- Polls Bridge `/health` for Desktop slot with `cdp_browser_running`
- Navigates to each 1688 offer URL via `/cdp/navigate`
- Extracts title, images, price, specs via `/cdp/evaluate`
- Writes enriched data to MongoDB `premiumproducts.products`
- Deploy via cron: `cronjob create no_agent=true script=1688_scrape_cdp.py schedule="every 5m"`
## Product Detail Enrichment Pipeline (Submit → Full Data)
For products submitted via Desktop Submit:
| Step | What | Tool |
|------|------|------|
| 1. Scrape | Extract raw data from 1688 page | CDP browser_console |
| 2. Structure | LLM extracts title/price/images/specs | Qwen 35B (:40006) / qwen3-vl |
| 3. Translate | zh → en, zh → ru | Qwen 35B / API |
| 4. Images | Remove watermark, resize 800×800 | SDXL/FLUX + AtomK toolbox |
| 5. Classify | Auto-categorize product | LLM |
| 6. Write | Update MongoDB document | Server endpoint |
## Reference Files
- `references/cdp-detail-extraction.md` — Extract structured product data from 1688 detail pages via Desktop CDP (body.innerText, image URL extraction, expression filter workarounds)
- `references/image-download-conversion.md` — Download alicdn product images and convert WebP→JPG in Cloud Bridge sessions (cron no_agent pattern)
- `references/mongodb-product-update.md` — Update product pool data via direct MongoDB connection when REST API is read-only (credentials, field schema, pitfalls)
## Scripts
- `scripts/1688_scrape_cdp.py` — Automated 1688 detail page scraper: health check → find Desktop slot → navigate → extract title/price/images/specs/weight → write MongoDB. Designed for cron `no_agent=true` polling: silently exits when no Desktop connected, runs when Desktop comes online.
## Pitfalls
- **Do NOT use search pages** (s.1688.com with keyword search forces login popup)
- **Category pages work better** — products visible without login
- **Product detail links** (offer/XXXX format) use encrypted redirects, hard to extract direct URLs
- **CamouFox on server IP = still blocked** — 1688 blocks at IP layer (datacenter detection), not browser fingerprint. CamouFox alone is insufficient; must pair with residential proxy or use Firecrawl
- **Free proxies are unstable** — may work one moment, blocked the next
- **Firecrawl waitFor=5000** is critical — without it, 1688's lazy-loaded content won't render
- **🔴 Firecrawl detail page TIMEOUT (2026-07-06 verified)**: Firecrawl API times out (120s+) on `detail.1688.com/offer/XXX` URLs. It works for category/search listing pages only. For individual detail pages, use Desktop CDP (Chinese IP).
- **🔴 AliExpress reference URLs are JS-only SPAs**: `aliexpress.com/item/XXX.html` returns a 2KB JS bootloader with no product data. The internal API (`aeglodetailweb/api/header`) also blocks non-browser requests with x5sec protection. AliExpress cannot be used as a fallback data source from server-side code.
- **🔴 Desktop must be connected for CDP scraping**: When Desktop is not running (0 slots on Bridge /health), CDP scraping is unavailable. Use a cron polling script (`no_agent=true`) that silently exits when no Desktop is connected, and runs the scraper when it comes online. See `scripts/1688_scrape_cdp.py` for the canonical pattern.
- **Firecrawl API key can expire** — returns `{"success": false, "error": "Unauthorized: Invalid token"}`. Fall back to Desktop CDP browser (Chinese IP) or renew the key.
- **Server-side CDP browser gets blocked** — `browser_navigate` to 1688 from SG3/SG5 returns `bixi.alicdn.com/punish/...` (Access Denied). Only Desktop Chrome on Chinese residential IP works for CDP-based scraping.
- **MOQ info** not reliably available from category pages; need detail page access for that
- **Prices shown may be tier prices** — the lowest price often requires high MOQ
- **Sales figures** on 1688 can be inflated; treat as directional, not exact
## Output Format
Save results as:
- `/tmp/1688_cross_border_top10.json` — Full enriched data
- `/tmp/1688_cross_border_top10.csv` — Spreadsheet-friendly format
CSV columns: 排名, 品类, 产品名称, 1688价格(元), 1688销量, 预估重量, TEMU参考价($), Amazon参考价($), 预估毛利率, 跨境标签, 推荐指数
## Related Skills
- `web-scraping-toolkit` — General scraping tools reference (Firecrawl is one of 5 tools documented there)
- `atomk-api-export` — Alternative product sourcing from AtomK system
- `temu-pipeline` — Downstream TEMU order processing