pdf-reading skill
Use this skill when you need to read, inspect, or extract content from PDF files — especially when file content is NOT in your context and you need to read it from disk. Covers content inventory, text extraction, page rasterization for visual inspection, embedded image/attachment/table/form-field extraction, and choosing the right reading strategy for different document types (text-heavy, scanned, slide-decks, forms, data-heavy). Do NOT use this skill for PDF creation, form filling, merging, splitting, watermarking, or encryption — use the pdf skill instead.
Is the pdf-reading skill safe?
Clean: nothing in its files matched our rules. We read 3 files in the folder on 2026-09-28.
No findings.
Install the pdf-reading skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/Razshy/Wiggle.git /tmp/Wiggle mkdir -p ~/.claude/skills cp -r /tmp/Wiggle/mnt-skills/public/pdf-reading ~/.claude/skills/pdf-reading
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
PDF Processing Guide
Overview
This guide covers essential PDF reading operations using Python libraries and command-line tools. For advanced features (pypdfium2 rendering, pdfplumber table settings, OCR fallback, encrypted/corrupted PDF handling), see REFERENCE.md.
Reading & Inspecting PDFs
Before doing anything with a PDF, understand what you're working with.
Content inventory
Run a quick diagnostic first. For simple tasks ("summarize this document"), pdfinfo + pdffonts + a text sample may suffice. For anything involving figures, attachments, or extraction issues, run the full set:
# Always: page count, file size, PDF version, metadata
pdfinfo document.pdf
# Always: does a text layer exist? No fonts → scanned/raster → see "Scanned documents"
pdffonts document.pdf
# If fonts are present: sample the text layer
pdftotext -f 1 -l 1 document.pdf - | head -20
# If figures/charts may matter:
pdfimages -list document.pdf
# If the PDF might contain embedded files (reports, portfolios):
pdfdetach -list document.pdfThis tells you:
means the PDF is scanned or raster-only: pdftotext will return nothing, so skip straight to "Scanned documents" below. Fonts shown as not embedded ("emb: no") with custom encodings may produce wrong characters on extraction.
- Page count and size — how big is the job?
- Font status — are any fonts present? An empty pdffonts table
clean text, or is it garbled (broken encoding)?
- Text extractability — when fonts exist, does pdftotext return
(Note: vector-drawn charts from matplotlib/Excel won't appear — see "Extracting embedded images" below)
- Embedded raster images — are there photos or raster figures?
- Attachments — are there embedded spreadsheets, data files, etc.?
Text extraction
pypdf for basic text:
from pypdf import PdfReader
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
# Extract text
text = ""
for page in reader.pages:
text += page.extract_text()pdftotext preserving layout (better for multi-column docs):
# Layout mode preserves spatial positioning
pdftotext -layout document.pdf output.txt
# Specific page range
pdftotext -f 1 -l 5 document.pdf output.txtpdfplumber for layout-aware extraction with positioning data:
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
for page in pdf.pages:
text = page.extract_text()
print(text)Visual inspection (rasterize pages)
Text extraction is blind to charts, diagrams, figures, equations, multi-column layout, and form structures. When any of these matter, rasterize the relevant page and Read the image:
# Rasterize a single page (page 3 here) at 150 DPI
pdftoppm -jpeg -r 150 -f 3 -l 3 document.pdf /tmp/page
# pdftoppm zero-pads the output filename based on TOTAL page count
# (e.g., page-03.jpg for a 50-page PDF, page-003.jpg for 200+ pages)
# Don't guess the filename — find it:
ls /tmp/page-*.jpgThen Read the resulting image file. This gives you full visual understanding of that page — layout, charts, equations, everything.
When to rasterize vs. text-extract:
for data, image for context — this is what Claude's API does natively with PDF uploads)
- Content/data questions → text extraction (cheaper, searchable)
- Figures, charts, visual layout → rasterize the page
- Tables → try text extraction first, rasterize if garbled
- Precision matters → do both (extract text AND rasterize; use text
Token cost awareness:
- Text extraction: ~200–400 tokens per page
- Rasterized image: ~1,600 tokens per page (at 150 DPI)
- Both together: ~2,000–2,400 tokens per page
For a 100-page PDF, rasterizing everything would consume ~160K tokens. Only rasterize pages that matter for the question at hand.
Choosing your reading strategy
Text-heavy documents (reports, articles, books): → Text extraction is primary. Rasterize only for specific figures or pages where layout matters.
Scanned documents (pdffonts shows no fonts): → pdftotext will return nothing — don't run it. Rasterize pages at 150 DPI and Read them visually. For bulk text extraction, use OCR (pytesseract after converting pages to images — see REFERENCE.md for a complete example).
Slide-deck PDFs (exported presentations): → Every page is primarily visual. Rasterize individual pages on demand. Text extraction gives you bullet-point text but loses all layout.
Form-heavy documents: → Extract form field values programmatically first (see below). Rasterize the form page for visual context if needed.
Data-heavy documents (tables, charts, figures): → Use pdfplumber for tables. Rasterize pages with charts/figures. Extract text for surrounding narrative. Consider both text AND image for the same page when precision matters.
Extracting embedded images
# List all embedded images with metadata (size, color, compression)
pdfimages -list document.pdf
# Extract all images as PNG
pdfimages -png document.pdf /tmp/img
# Extract from specific pages only (pages 3-5)
pdfimages -png -f 3 -l 5 document.pdf /tmp/img
# Extract in original format (JPEG stays JPEG, etc.)
pdfimages -all document.pdf /tmp/imgThen Read /tmp/img-000.png (etc.) to see each extracted image.
Gotcha — vector graphics: pdfimages extracts only raster image data. Charts and diagrams drawn as vector graphics (common in matplotlib, Excel, and R exports) will NOT appear — they are page content operators, not image objects. For these, rasterize the whole page with pdftoppm instead.
Gotcha — empty images: pdfimages sometimes produces many tiny or empty image files — these are typically background masks, transparency layers, or decorative elements. Filter by file size to find the real content images.
Programmatic extraction with position data:
import fitz # PyMuPDF
doc = fitz.open("document.pdf")
for page in doc:
for img in page.get_images():
xref = img[0]
pix = fitz.Pixmap(doc, xref)
if pix.n - pix.alpha > 3: # CMYK or other non-RGB
pix = fitz.Pixmap(fitz.csRGB, pix)
pix.save(f"/tmp/img_{xref}.png")Extracting file attachments
PDFs can contain embedded files — spreadsheets, data files, other documents. Common in business reports, PDF portfolios, and PDF/A-3 compliance documents.
# List all attachments
pdfdetach -list document.pdf
# Extract all attachments to a directory
mkdir -p /tmp/attachments
pdfdetach -saveall -o /tmp/attachments/ document.pdf
# Extract a specific attachment by number (1-based index from -list output)
pdfdetach -save 1 -o /tmp/attachment.pdf document.pdfIn Python:
import os
from pypdf import PdfReader
reader = PdfReader("document.pdf")
for name, content_list in reader.attachments.items():
safe_name = os.path.basename(name) # sanitize — name comes from the PDF
for content in content_list:
with open(f"/tmp/{safe_name}", "wb") as f:
f.write(content)Two attachment mechanisms exist in PDFs: page-level file annotation attachments (shown as paperclip icons in viewers) and document-level embedded files (in the EmbeddedFiles name tree). Both pdfdetach and pypdf handle the common cases. Rich media assets (3D, video) embedded as annotations may not appear in the attachment list — use PyMuPDF to iterate page annotations for those.
Extracting form field data
PDFs with interactive forms (government forms, applications, contracts) have fillable fields whose values can be read programmatically:
from pypdf import PdfReader
reader = PdfReader("form.pdf")
# Text input fields only:
fields = reader.get_form_text_fields()
for name, value in fields.items():
print(f"{name}: {value}")
# All field types (checkboxes, radio buttons, dropdowns too):
all_fields = reader.get_fields() or {}
for name, field in all_fields.items():
print(f"{name}: {field.get('/V', '')} (type: {field.get('/FT', '')})")getformtextfields() returns only text input fields. For government forms and contracts that use checkboxes, radio buttons, and dropdowns, use getfields() instead to see all field types.
For comprehensive field info (types, options, defaults):
pdftk form.pdf dump_data_fieldsFor anything beyond reading form data — filling forms, creating forms — use the pdf skill — invoke it by name if you have a Skill tool, or Read its SKILL.md (listed in your available skills, or in the same skills directory as this file).
Audio, video, and other rare embedded content
More skills from Razshy/Wiggle
- Aalgorithmic-artCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
- Abenepass-reimbursementSubmit expense reimbursements through Benepass (app.getbenepass.com). For users whose employer uses Benepass as their benefits platform. Handles login, benefit selection, form filling, receipt upload, and submission. Requires browser/computer-use capabilities.
- Abrand-guidelinesApplies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.
- Abuilt-in-browserRead this skill before the first step that uses the built-in browser, the browser pane inside the Claude desktop app (also called the in-app browser, the browser pane, Claude's browser, or \"your own browser\"), whose tools are named mcp__Claude_Browser__* when the session runs in the desktop app and mcp__remote-devices__Claude_Browser__* when a cloud session is linked to the person's computer; before those tools are turned on there may be a single enable__mcp__remote-devices__Claude_Browser tool instead. It covers the pane's persistent sign-ins, tabs and preview_start, reading pages as text, site approvals, what the pane cannot open, and what to do when it cannot be reached. It is not for Claude in Chrome (mcp__claude-in-chrome__* tools), which has its own skill, and it does not decide which browser to use.
- Acall-to-bookMake a phone call to book an appointment or reservation. Checks calendar first, gets explicit consent before dialing, discloses AI identity on the call, and adds the booking to calendar when done.
- Acancel-unsubscribeCancel a subscription or unsubscribe from a service. Works from a description, a pasted charge line, a URL, or a photo/screenshot. Can also audit a full statement for recurring charges and cancel several at once. Finds the right contact method and handles the cancellation — including phone calls.
- Acanvas-designCreate beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
- Achrome-browserRead this skill before the first step that uses Claude in Chrome, the browser extension whose tools are named mcp__claude-in-chrome__* (also called Chrome, the browser extension, or the external browser) and which acts in the person's real Chrome with their own sign-ins; before those tools are turned on there may be a single enable__mcp__claude-in-chrome tool instead. It covers loading the tools in one ToolSearch call, checking the person's open tabs and working in a new tab, site permissions, GIF recordings, console logs, dialogs to avoid, and when to stop and ask. It is not for the built-in browser (mcp__Claude_Browser__* or mcp__remote-devices__Claude_Browser__* tools), which has its own skill, and it does not decide which browser to use.
- Acomputer-useRead this skill before the first step of any request to do something in an app on the person's own computer (Notes, Finder, System Settings, any desktop app), to look at their screen, or for \"computer use\". Computer use (desktop control) lets Claude take screenshots of the person's desktop and control it with clicks, typing and scrolling through the Claude desktop app; its tools are named mcp__computer-use__* when the session runs in the desktop app and mcp__remote-devices__computer_* when a cloud session is linked to the person's computer; before a conversation is linked there may be no such tools, only an enable__mcp__remote-devices__Claude_Browser tool, which links it. It covers getting linked, picking the right tool, the access flow, and what to do when computer use is off or out of reach. It is not for websites, which go through Claude in Chrome or the built-in browser and their own skills.
- Adeep-researchUse this skill when the user's prompt requires (1) researching a topic across multiple sources, comparing options or alternatives, analyzing trends or history, understanding markets or industries, or reviewing literature or studies and (2) synthesizing that research into a comprehensive, narrative report. If you're planning to search the web or internal knowledge bases, consider using this skill. This skill coordinates research subagents, so use it only when you have a tool for spawning subagents (the Agent or Task tool); otherwise, research the question directly.
- Adoc-coauthoringGuide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.
- AdocxUse this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx) or Word templates (.dotx). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file (to download, email or print), use this skill. However, if they ask for a document, page, report, memo, or notes WITHOUT naming a file format and the session offers a dedicated document or page skill or connector, use that instead. Do NOT use for PDFs, spreadsheets, Google Docs, or coding unrelated to document generation.