Files

3.5 KiB

name, description, tags
name description tags
drissionpage-website-analysis Use DrissionPage to comprehensively analyze websites — structure, content, links, images, forms, scripts, SEO, and performance metrics. Creates a detailed report with screenshots.
web-analysis
drissionpage
automation
scraping
seo
python

DrissionPage Website Analysis

Use DrissionPage to perform comprehensive website analysis — structure, content, links, images, forms, scripts, SEO, and performance.

Setup

# Install DrissionPage (use a temp venv if needed)
python3 -m venv /tmp/drission_env && source /tmp/drission_env/bin/activate
pip install DrissionPage

Browser Configuration

from DrissionPage import ChromiumPage, ChromiumOptions

co = ChromiumOptions()
co.set_browser_path('/home/ubuntu/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome')
co.headless()
co.set_argument('--no-sandbox')
co.set_argument('--disable-dev-shm-usage')
co.set_argument('--disable-gpu')
co.auto_port()
page = ChromiumPage(co)

Analysis Categories

A comprehensive website analysis should cover:

  1. Page Basic Info — URL, title, language, charset, cookies, storage
  2. Page Structure — element counts, tag distribution
  3. Content Extraction — headings (H1-H6), paragraphs, meta tags
  4. Link Analysis — internal/external/mail/JS links
  5. Image Analysis — count, alt attributes, lazy loading, dimensions
  6. Form Analysis — form count, actions, methods, input fields
  7. Scripts & Styles — external vs inline JS, CSS stylesheets
  8. SEO Analysis — meta description, keywords, robots, canonical, Open Graph, Twitter Card, H1 count
  9. Performance — DNS, TCP, TTFB, DOM ready, full load time (via JS performance.timing)
  10. Screenshot — full page screenshot for visual verification

Key Code Patterns

Get all elements of a type

links = page.eles('tag:a')
images = page.eles('tag:img')
headings = page.eles('tag:h1')

Run JavaScript on the page

info = page.run_js('return { title: document.title, url: window.location.href };')

Get element attributes

for link in links:
    href = link.attr('href')
    text = link.text.strip()

Take screenshot

page.get_screenshot(path='/tmp/screenshot.png', full_page=True)

Performance timing

perf = page.run_js('''
return {
    loadTime: performance.timing.loadEventEnd - performance.timing.navigationStart,
    domReady: performance.timing.domContentLoadedEventEnd - performance.timing.navigationStart,
    ttfb: performance.timing.responseStart - performance.timing.navigationStart,
};
''')

Pitfalls

  • uv-managed Python: The environment may use uv-managed Python which doesn't allow direct pip install. Create a temp venv instead.
  • Chrome path: The Playwright-installed Chromium is at /home/ubuntu/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome.
  • Headless mode: Always use co.headless() for server environments.
  • Auto port: Use co.auto_port() to avoid port conflicts with existing browser instances.
  • Element limits: When counting all elements (page.eles('xpath://*')), limit to ~5000 to avoid performance issues.
  • Wait for content: Add time.sleep(3) after page load for JS-rendered content to appear.

Output Structure

Generate a structured report with:

  • Summary table (site name, tech stack, language, load time)
  • Section-by-section analysis with emoji headers
  • SEO issues flagged with or ⚠️
  • Full page screenshot for visual reference