web-scraping skill
Extract data from websites, including JavaScript-rendered SPAs and dynamic content
Is the web-scraping skill safe?
Read the findings before you install it. We read 13 files in the folder on 2026-09-28.
- high
references/security-audit-pattern.md:35Reads credential files (SSH keys, cloud or package-manager tokens) that a skill has no normal reason to touch.
("/api/../../../etc/passwd", "GET", None),
Install the web-scraping skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/kevinnft/ai-agent-skills.git /tmp/ai-agent-skills mkdir -p ~/.claude/skills cp -r /tmp/ai-agent-skills/skills/research/web-scraping ~/.claude/skills/web-scraping
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Web Scraping
Extract structured data from websites, handling both static HTML and JavaScript-rendered content (React, Next.js, Vue, etc.).
When to Use
- User asks to "scrape", "extract", or "get data from" a website
- Target site uses client-side rendering (SPA frameworks)
- Need to interact with dynamic content (infinite scroll, lazy loading)
- API endpoints are not available or documented
Approach Selection
1. Static HTML (curl + parsing)
Use when: Site serves complete HTML without JavaScript rendering.
curl -sL 'https://example.com' | grep -oP 'pattern'
# or with jq for JSON APIs
curl -s 'https://api.example.com/data' | jq '.items[]'Pros: Fast, lightweight, no dependencies Cons: Fails on JS-rendered content
2. Headless Browser (Puppeteer/Playwright)
Use when: Content is rendered client-side (React, Next.js, Vue, Angular).
Node.js + Puppeteer (recommended for WSL2/containers):
const puppeteer = require('puppeteer');
const browser = await puppeteer.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox'] // Required in WSL2/containers
});
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'networkidle2',
timeout: 60000
});
// Wait for dynamic content
await new Promise(resolve => setTimeout(resolve, 3000));
// Extract text
const content = await page.evaluate(() => document.body.innerText);
// Extract structured data
const data = await page.evaluate(() => {
return Array.from(document.querySelectorAll('.item')).map(el => ({
title: el.querySelector('.title')?.innerText,
value: el.querySelector('.value')?.innerText
}));
});
await browser.close();Pros: Handles all JS rendering, can interact with page Cons: Slower, heavier resource usage
3. API Inspection (DevTools Network tab)
Use when: Site loads data via XHR/fetch calls.
- Open browser DevTools → Network tab
- Filter by XHR/Fetch
- Find API endpoint
- Replicate with curl/fetch
Pros: Fastest, most reliable Cons: Requires manual inspection, may need auth tokens
Alternative: Reverse-engineer from minified JS (when browser access blocked):
Method A: Direct curl (if no Cloudflare)
# Download main JS bundle
curl -s "https://example.com/assets/main-[hash].js" > /tmp/bundle.js
# Search for API patterns
grep -oP '"/[a-z_/-]{3,}"' /tmp/bundle.js | sort -u
strings /tmp/bundle.js | grep -i 'keyword' | head -20Method B: TinyFish browser automation (if Cloudflare protected)
When curl fails due to Cloudflare Turnstile, use TinyFish to bypass protection and download JS via Chrome DevTools Protocol:
# 1. Create TinyFish browser session (bypasses Cloudflare)
# 2. Wait for challenge completion
# 3. Connect to browser via CDP WebSocket
# 4. Use Runtime.evaluate to fetch JS files
# 5. Extract API endpoints from minified codeSee references/tinyfish-js-reverse-engineering.md for full workflow (tested on rpow2swap.com May 2026).
Trial-error common paths with size check:
for path in /api/listings /api/orders /listings /tokens /api/stats; do
echo "Testing: https://example.com$path"
timeout 3 curl -s -m 3 -o /dev/null -w "HTTP %{http_code} | Size: %{size_download} bytes\\n" \
"https://example.com$path" 2>&1 || echo "Timeout/Error"
done
# Look for large responses (>10KB = likely data endpoint, <2KB = likely SPA HTML)Success indicators:
- Response size >10KB → likely JSON data endpoint
- Response size <2KB → likely SPA HTML fallback
- Timeout → endpoint exists but slow/protected
See references/spa-api-discovery.md for full technique (tested on rpow2swap.com May 2026).
Success case (rpow2swap.com, May 2026):
# 1. Try common API paths with timeout
for path in /api/listings /api/tokens /listings /tokens /api/orderbook; do
timeout 3 curl -s -m 3 -o /dev/null -w "HTTP %{http_code} | Size: %{size_download}\\n" \
"https://example.com$path"
done
# Result: /api/listings returned 65KB (200 OK) — found it!
# 2. Fetch and inspect data
curl -s "https://example.com/api/listings" | head -c 2000
# Returns JSON array with full listing data
# 3. Build monitoring bot
# State-based change detection: track seen IDs, alert on new entriesKey insight: Many SPAs use predictable REST paths (/api/). Trial-error with timeout is faster than reverse-engineering minified JS.
3.5. Third-Party APIs (Twitter/X)
Use when: Scraping Twitter/X content (tweets, profiles, media).
Primary: vxtwitter API (no auth, works from terminal)
# Get tweet data
curl -s "https://api.vxtwitter.com/Twitter/status/{tweet_id}" | jq -r '.tweet | {text, author, likes, retweets, replies, media}'
# Get account info
curl -s "https://api.vxtwitter.com/{handle}" | jq -r '.user | {name, description, followers, website}'
# Extract quoted tweet (QRT)
curl -s "https://api.vxtwitter.com/Twitter/status/{tweet_id}" | jq -r '.qrt | {text, author, likes}'Fallback: fxtwitter API (same structure)
curl -s "https://api.fxtwitter.com/{handle}/status/{tweet_id}"Pros: No auth, fast, structured JSON, includes media URLs Cons: Rate limited, may lag behind real-time data
Note: Twitter's official API requires auth and has strict rate limits. Use vxtwitter/fxtwitter for read-only access.
4. Cloud Browser Services (Cloudflare bypass)
Use when: Site has Cloudflare Turnstile, bot detection, or anti-scraping measures.
Browserbase (recommended, tested May 2026):
import requests
# Create session
response = requests.post(
"https://www.browserbase.com/v1/sessions",
headers={"X-BB-API-Key": API_KEY, "Content-Type": "application/json"},
json={"projectId": PROJECT_ID}
)
session = response.json()
debug_url = f"https://www.browserbase.com/v1/sessions/{session['id']}/debug"See references/browserbase-api.md for full API reference.
Browser Use Cloud (alternative, had 404 issues May 2026):
response = requests.post(
"https://api.browser-use.com/api/v1/run-task",
headers={"Authorization": f"Bearer {API_KEY}"},
json={"task": "Go to protected-site.com and extract data"}
)See references/browser-use-cloud-api.md for troubleshooting.
Pros: Bypasses Cloudflare, residential proxies, no local resources Cons: Paid service, requires API key, SDK needed for automation
WSL2 / Container Considerations
Chrome sandbox issues are common in WSL2 and Docker. Always use:
args: ['--no-sandbox', '--disable-setuid-sandbox']Python venv issues: WSL2 Ubuntu may lack python3-venv. Use Node.js approach instead or install:
sudo apt install python3.12-venvWorkflow
- Try curl first — check if content is in initial HTML
- Inspect Network tab — look for API endpoints
- Use headless browser — if content is JS-rendered
- Extract incrementally — get raw text first, then refine selectors
Pitfalls
- Don't assume static HTML — modern sites often use SSR/CSR hybrid (Next.js)
- Wait for content — add delays after page load for dynamic content
- Check robots.txt — respect crawling policies
- Rate limiting — add delays between requests for bulk scraping
- User-Agent — some sites block default headless browser UA
- Cloudflare protection — sites with Cloudflare Turnstile/challenge pages block curl and standard browsers. Use Browser Use Cloud or stealth browser libraries.
- VPS browser limitations — Hermes browser tool may fail on VPS with sandbox errors. Use --no-sandbox flag or cloud browser services.
- Browser Use API confusion — Browser Use has TWO APIs: open-source library (local, free) vs Cloud API (managed, paid). Cloud API endpoint structure is confusing (examples use /api/v1/, docs say /v3/). If getting 404 errors, see references/browser-use-cloud-api.md for troubleshooting.
Verification
More skills from kevinnft/ai-agent-skills
- Aaddyosmani-tddDrives development with tests. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.
- AairtableAirtable REST API via curl. Records CRUD, filters, upserts.
- Aapi-and-interface-designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
- Aapi-monitoring-botsBuild monitoring bots that poll APIs and send notifications on state changes (new listings, price alerts, status updates)
- Aapple-notesManage Apple Notes via memo CLI: create, search, edit.
- Aapple-remindersApple Reminders via remindctl: add, list, complete.
- Aarchitecture-diagramDark-themed SVG architecture/cloud/infra diagrams as HTML.
- AarxivSearch arXiv papers by keyword, author, category, or ID.
- Aascii-artASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
- Aascii-videoASCII video: convert video/audio to colored ASCII MP4/GIF.
- AaudiocraftAudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
- CaxolotlAxolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).