Files

101 lines
3.5 KiB
Markdown

---
name: drissionpage-website-analysis
description: Use DrissionPage to comprehensively analyze websites — structure, content, links, images, forms, scripts, SEO, and performance metrics. Creates a detailed report with screenshots.
tags: [web-analysis, drissionpage, automation, scraping, seo, python]
---
# DrissionPage Website Analysis
Use DrissionPage to perform comprehensive website analysis — structure, content, links, images, forms, scripts, SEO, and performance.
## Setup
```bash
# Install DrissionPage (use a temp venv if needed)
python3 -m venv /tmp/drission_env && source /tmp/drission_env/bin/activate
pip install DrissionPage
```
## Browser Configuration
```python
from DrissionPage import ChromiumPage, ChromiumOptions
co = ChromiumOptions()
co.set_browser_path('/home/ubuntu/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome')
co.headless()
co.set_argument('--no-sandbox')
co.set_argument('--disable-dev-shm-usage')
co.set_argument('--disable-gpu')
co.auto_port()
page = ChromiumPage(co)
```
## Analysis Categories
A comprehensive website analysis should cover:
1. **Page Basic Info** — URL, title, language, charset, cookies, storage
2. **Page Structure** — element counts, tag distribution
3. **Content Extraction** — headings (H1-H6), paragraphs, meta tags
4. **Link Analysis** — internal/external/mail/JS links
5. **Image Analysis** — count, alt attributes, lazy loading, dimensions
6. **Form Analysis** — form count, actions, methods, input fields
7. **Scripts & Styles** — external vs inline JS, CSS stylesheets
8. **SEO Analysis** — meta description, keywords, robots, canonical, Open Graph, Twitter Card, H1 count
9. **Performance** — DNS, TCP, TTFB, DOM ready, full load time (via JS `performance.timing`)
10. **Screenshot** — full page screenshot for visual verification
## Key Code Patterns
### Get all elements of a type
```python
links = page.eles('tag:a')
images = page.eles('tag:img')
headings = page.eles('tag:h1')
```
### Run JavaScript on the page
```python
info = page.run_js('return { title: document.title, url: window.location.href };')
```
### Get element attributes
```python
for link in links:
href = link.attr('href')
text = link.text.strip()
```
### Take screenshot
```python
page.get_screenshot(path='/tmp/screenshot.png', full_page=True)
```
### Performance timing
```python
perf = page.run_js('''
return {
loadTime: performance.timing.loadEventEnd - performance.timing.navigationStart,
domReady: performance.timing.domContentLoadedEventEnd - performance.timing.navigationStart,
ttfb: performance.timing.responseStart - performance.timing.navigationStart,
};
''')
```
## Pitfalls
- **uv-managed Python**: The environment may use uv-managed Python which doesn't allow direct pip install. Create a temp venv instead.
- **Chrome path**: The Playwright-installed Chromium is at `/home/ubuntu/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome`.
- **Headless mode**: Always use `co.headless()` for server environments.
- **Auto port**: Use `co.auto_port()` to avoid port conflicts with existing browser instances.
- **Element limits**: When counting all elements (`page.eles('xpath://*')`), limit to ~5000 to avoid performance issues.
- **Wait for content**: Add `time.sleep(3)` after page load for JS-rendered content to appear.
## Output Structure
Generate a structured report with:
- Summary table (site name, tech stack, language, load time)
- Section-by-section analysis with emoji headers
- SEO issues flagged with ❌ or ⚠️
- Full page screenshot for visual reference