Add archived/cdp-ozon-collector
This commit is contained in:
@@ -0,0 +1,109 @@
|
|||||||
|
---
|
||||||
|
name: cdp-ozon-collector
|
||||||
|
description: Auto-surf Ozon.ru via CDP tunnel — randomly hops between product pages, extracts product data (SKU, price, rating, seller, stock, delivery), saves to JSONL. Runs continuously until killed.
|
||||||
|
version: 1.0
|
||||||
|
tags: [ozon, cdp, collector, scraping, auto-surf]
|
||||||
|
---
|
||||||
|
|
||||||
|
# CDP Ozon Auto-Surfer Collector
|
||||||
|
|
||||||
|
Uses the Cloud Bridge CDP tunnel to control a real Chrome browser open on Ozon.ru, automatically hopping between product pages and extracting data.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
Agent → HTTP :9229 → WS :9228 → Desktop Electron → Extension WS → chrome.debugger → Ozon page
|
||||||
|
```
|
||||||
|
|
||||||
|
- **API Base**: `http://127.0.0.1:9229`
|
||||||
|
- **API Key**: `Bing2026Cao$$$`
|
||||||
|
- **Key endpoints**: `/cdp/attach`, `/cdp/detach`, `/cdp/evaluate`, `/cdp/navigate`, `/cdp/send`
|
||||||
|
|
||||||
|
## Script Location
|
||||||
|
|
||||||
|
`/home/ubuntu/ozon_collector.py`
|
||||||
|
|
||||||
|
## How It Works
|
||||||
|
|
||||||
|
1. Finds Ozon page targets via `Target.getTargets`
|
||||||
|
2. Attaches via CDP, extracts product data from DOM
|
||||||
|
3. Finds all `/product/` links on the current page
|
||||||
|
4. Picks a random **unvisited** link and navigates to it
|
||||||
|
5. If no fresh links → searches a random keyword from a built-in list
|
||||||
|
6. Repeats with 30-90s random delays between cycles
|
||||||
|
|
||||||
|
## Data Extraction (JS expression in `collect_page()`)
|
||||||
|
|
||||||
|
Ozon uses obfuscated CSS classes (`.a9y6`, `.kp6`, etc.) that change frequently. **Do NOT rely on class selectors.** Instead:
|
||||||
|
|
||||||
|
- **Price**: Parse `document.body.innerText` lines containing `₽` (U+8381). Skip promo line "Товары за 1₽", take 2nd match as card price.
|
||||||
|
- **Rating**: Regex `(\d+\.\d+)\s*[•·]\s*(\d+\s*отзыв)` from body text
|
||||||
|
- **Seller**: Regex `Магазин[\s\S]*?\n([\w\s]+?)\n.*О магазине` or fallback to `a[href*="seller"]`
|
||||||
|
- **SKU**: Extract from URL pattern `/product/...-(\d+)/`
|
||||||
|
- **Stock**: Regex `(\d+[\s\d]*\s*(?:шт|единиц)\s*осталось)`
|
||||||
|
- **Images**: Filter `img[src*="cdn"]` containing "ozon" or "wbcdn", exclude "payments-cdn" and "marketing-api"
|
||||||
|
- **Delivery**: Regex `Доставим\s*(с\s*\d+[^\n]*)`
|
||||||
|
|
||||||
|
## Output Format
|
||||||
|
|
||||||
|
JSONL file at `/home/ubuntu/ozon_collection.jsonl`:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"timestamp": "2026-05-19T06:51:36.123456",
|
||||||
|
"cycle": 42,
|
||||||
|
"run_id": "20260519_012342",
|
||||||
|
"title": "...",
|
||||||
|
"url": "https://www.ozon.ru/product/...",
|
||||||
|
"h1": "Product name",
|
||||||
|
"price": "85 ₽",
|
||||||
|
"allPrices": ["Товары за 1₽", "85 ₽", "86 ₽", "2 146 ₽"],
|
||||||
|
"rating": "4.8 • 746 отзыв",
|
||||||
|
"ratingValue": "4.8",
|
||||||
|
"reviewCount": "746 отзыв",
|
||||||
|
"seller": "",
|
||||||
|
"sku": "4267896148",
|
||||||
|
"imgs": ["https://..."],
|
||||||
|
"stock": "99 единиц осталось",
|
||||||
|
"delivery": "с 4 июня"
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Pitfalls & Lessons Learned
|
||||||
|
|
||||||
|
1. **No class selectors for Ozon** — They obfuscate and rotate CSS classes. Always use text-based regex parsing from `body.innerText`.
|
||||||
|
2. **Price parsing order matters** — First ₽ line is often promo "Товары за 1₽". Real price is 2nd+ line.
|
||||||
|
3. **Duplicate tabs** — `Target.getTargets` returns both the Ozon homepage and product sub-tab. Deduplicate by URL (strip query params).
|
||||||
|
4. **CDP attach follows openerId** — When attaching to a child tab, CDP may return the parent's targetId. Use `/cdp/navigate` to go directly to the desired URL instead.
|
||||||
|
5. **Page load timing** — After navigate, wait 3-6 seconds (randomized) before extracting data. Too short → empty h1/price.
|
||||||
|
6. **Seller extraction unreliable** — Ozon renders seller info via React hydration; regex may miss it. The `a[href*="seller"]` fallback also often fails.
|
||||||
|
7. **Images need filtering** — Raw `img` selector picks up payment icons and marketing banners. Filter out `payments-cdn` and `marketing-api` URLs.
|
||||||
|
8. **Python output buffering** — Run with `python3 -u` for unbuffered output in background processes.
|
||||||
|
|
||||||
|
## Running
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Start
|
||||||
|
python3 -u /home/ubuntu/ozon_collector.py &
|
||||||
|
|
||||||
|
# Check status
|
||||||
|
ps aux | grep ozon_collector | grep -v grep
|
||||||
|
wc -l /home/ubuntu/ozon_collection.jsonl
|
||||||
|
|
||||||
|
# Stop
|
||||||
|
pkill -f ozon_collector.py
|
||||||
|
```
|
||||||
|
|
||||||
|
## Monitoring Queries
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Count unique SKUs
|
||||||
|
grep -oP '"sku": "\K[^"]+' /home/ubuntu/ozon_collection.jsonl | sort -u | wc -l
|
||||||
|
|
||||||
|
# Latest 5 records
|
||||||
|
tail -5 /home/ubuntu/ozon_collection.jsonl | python3 -c "
|
||||||
|
import sys,json
|
||||||
|
for l in sys.stdin:
|
||||||
|
d=json.loads(l)
|
||||||
|
print(f\"{d['timestamp'][:19]} | SKU={d.get('sku','')} | {d.get('h1','')[:55]} | {d.get('price','')} | ⭐{d.get('ratingValue','')}\")"
|
||||||
|
```
|
||||||
Reference in New Issue
Block a user