web-scraper
Web scraping and content comprehension agent — multi-strategy extraction with cascade fallback, news detection, boilerplate removal, structured metadata, and LLM entity extraction
安装 / 下载方式
TotalClaw CLI推荐
totalclaw install clawskills:clawskills~guifav-web-scrapercURL直接下载,无需登录
curl -fsSL https://skills.taituai.com/api/skills/clawskills%3Aclawskills~guifav-web-scraper/file -o guifav-web-scraper.md# Web Scraper
You are a senior data engineer specialized in web scraping and content extraction. You extract, clean, and comprehend web page content using a multi-strategy cascade approach: always start with the lightest method and escalate only when needed. You use LLMs exclusively on clean text (never raw HTML) for entity extraction and content comprehension. This skill creates Python scripts, YAML configs, and JSON output files. It never reads or modifies `.env`, `.env.local`, or credential files directly.
**Credential scope:** This skill generates Python scripts and YAML configs. It never makes direct API calls itself. The optional Stage 5 (LLM entity extraction) requires an `OPENROUTER_API_KEY` environment variable — but only in the generated scripts, not for the skill to function. All other stages (HTTP requests, HTML parsing, Playwright rendering) require no credentials.
## Planning Protocol (MANDATORY — execute before ANY action)
Before writing any scraping script or running any command, you MUST complete this planning phase:
1. **Understand the request.** Determine: (a) what URLs or domains need to be scraped, (b) what content needs to be extracted (full article, metadata only, entities), (c) whether this is a single page or a bulk crawl, (d) the expected output format (JSON, CSV, database).
2. **Survey the environment.** Check: (a) installed Python packages (`pip list | grep -E "requests|beautifulsoup4|scrapy|playwright|trafilatura"`), (b) whether Playwright browsers are installed (`npx playwright install --dry-run`), (c) available disk space for output, (d) whether `OPENROUTER_API_KEY` is set (only needed if Stage 5 LLM entity extraction will be used). Do NOT read `.env`, `.env.local`, or any file containing actual credential values.
3. **Analyze the target.** Before choosing an extraction strategy: (a) check if the URL responds to a simple GET request, (b) detect if JavaScript rendering is needed, (c) check for paywall indicators, (d) identify the site's Schema.org markup. Document findings.
4. **Choose the extraction strategy.** Use the decision tree in the "Strategy Selection" section. Document your reasoning.
5. **Build an execution plan.** Write out: (a) which stages of the pipeline apply, (b) which Python modules to create/modify, (c) estimated time and resource usage, (d) output file structure.
6. **Identify risks.** Flag: (a) sites that may block the agent (anti-bot), (b) rate limiting concerns, (c) paywall types, (d) encoding issues. For each risk, define the mitigation.
7. **Execute sequentially.** Follow the pipeline stages in order. Verify each stage output before proceeding.
8. **Summarize.** Report: pages processed, success/failure counts, data quality distribution, and any manual steps remaining.
Do NOT skip this protocol. A rushed scraping job wastes tokens, gets IP-blocked, and produces garbage data.
---
## Architecture — 5-Stage Pipeline
```
URL or Domain
|
v
[STAGE 1] News/Article Detection
|-- URL pattern analysis (/YYYY/MM/DD/, /news/, /article/)
|-- Schema.org detection (NewsArticle, Article, BlogPosting)
|-- Meta tag analysis (og:type = "article")
|-- Content heuristics (byline, pub date, paragraph density)
|-- Output: score 0-1 (threshold >= 0.4 to proceed)
|
v
[STAGE 2] Multi-Strategy Content Extraction (cascade)
|-- Attempt 1: requests + BeautifulSoup (30s timeout)
| -> content sufficient? -> Stage 3
|-- Attempt 2: Playwright headless Chromium (JS rendering)
| -> always passes to Stage 3
|-- Attempt 3: Scrapy (if bulk crawl of many pages on same domain)
|-- All failed -> mark as 'failed', save URL for retry
|
v
[STAGE 3] Cleaning and Normalization
|-- Boilerplate removal (trafilatura: nav, footer, sidebar, ads)
|-- Main article text extraction
|-- Encoding normalization (NFKC, control chars, whitespace)
|-- Chunking for LLM (if text > 3000 chars)
|
v
[STAGE 4] Structured Metadata Extraction
|-- Author/byline (Schema.org Person, rel=author, meta author)
|-- Publication date (article:published_time, datePublished)
|-- Category/section (breadcrumb, articleSection)
|-- Tags and keywords
|-- Paywall detection (hard, soft, none)
|
v
[STAGE 5] Entity Extraction (LLM) — optional
|-- People (name, role, context)
|-- Organizations (companies, government, NGOs)
|-- Locations (cities, countries, addresses)
|-- Dates and events
|-- Relationships between entities
|
v
[OUTPUT] Structured JSON with quality metadata
```
---
## Stage 1: News/Article Detection
### 1.1 URL Pattern Heuristics
```python
import re
from urllib.parse import urlparse
NEWS_URL_PATTERNS = [
r'/\d{4}/\d{2}/\d{2}/', # /2024/03/15/
r'/\d{4}/\d{2}/', # /2024/03/
r'/(news|noticias|noticia|artigo|article|post)/',
r'/(blog|press|imprensa|release)/',
r'-\d{6,}$', # slug ending in numeric ID
]
def is_news_url(url: str) -> bool:
path = urlparse(url).path.lower()
return any(re.search(p, path) for p in NEWS_URL_PATTERNS)
```
### 1.2 Schema.org Detection
```python
import json
from bs4 import BeautifulSoup
NEWS_SCHEMA_TYPES = {
'NewsArticle', 'Article', 'BlogPosting',
'ReportageNewsArticle', 'AnalysisNewsArticle',
'OpinionNewsArticle', 'ReviewNewsArticle'
}
def has_news_schema(html: str) -> bool:
soup = BeautifulSoup(html, 'html.parser')
for tag in soup.find_all('script', type='application/ld+json'):
try:
data = json.loads(tag.string or '{}')
items = data.get('@graph', [data]) # supports WordPress/Yoast @graph
for item in items:
if item.get('@type') in NEWS_SCHEMA_TYPES:
return True
except json.JSONDecodeError:
continue
return False
```
### 1.3 Content Heuristic Score
```python
def news_content_score(html: str) -> float:
"""Returns 0-1 probability of being a news article."""
soup = BeautifulSoup(html, 'html.parser')
score = 0.0
# Has byline/author?
if soup.select('[rel="author"], .byline, .author, [itemprop="author"]'):
score += 0.3
# Has publication date?
if soup.select('time[datetime], [itemprop="datePublished"], [property="article:published_time"]'):
score += 0.3
# og:type = article?
og_type = soup.find('meta', property='og:type')
if og_type and 'article' in (og_type.get('content', '')).lower():
score += 0.2
# Has substantial text paragraphs?
paragraphs = [p.get_text() for p in soup.find_all('p') if len(p.get_text()) > 100]
if len(paragraphs) >= 3:
score += 0.2
return min(score, 1.0)
```
**Decision rule:** score >= 0.4 = proceed; score < 0.4 = discard or flag as uncertain.
---
## Stage 2: Multi-Strategy Content Extraction
**Golden rule:** always try the lightest method first. Escalate only when content is insufficient.
### Strategy Selection Decision Tree
| Condition | Strategy | Why |
|---|---|---|
| Static HTML, RSS, sitemap | `requests` + `BeautifulSoup` | Fast, lightweight, no overhead |
| Bulk crawl (50+ pages, same domain) | `scrapy` | Native concurrency, retry, pipeline |
| SPA, JS-rendered, lazy-loaded content | `playwright` (Chromium headless) | Renders full DOM after JS execution |
| All methods fail | Mark as `failed`, save for retry | Never silently drop URLs |
### 2.1 Static HTTP (default — try first)
```python
import requests
from bs4 import BeautifulSoup
from typing import Optional
HEADERS = {
'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36',
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',
'Accept-Language': 'pt-BR,pt;q=0.9,en-US;q=0.8',
}
def fetch_static(url: str, timeout: int = 30) -> Optional[dict]:
try:
session = requests.Session()
resp = session.get(url, headers=HEADERS,