citation-management skill
Comprehensive citation management for academic research. Search OpenAlex, PubMed, and Google Scholar for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
Is the citation-management skill safe?
Clean: nothing in its files matched our rules. We read 20 files in the folder on 2026-09-28.
No findings.
Install the citation-management skill
A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.
git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills.git /tmp/scientific-agent-skills mkdir -p ~/.claude/skills cp -r /tmp/scientific-agent-skills/skills/citation-management ~/.claude/skills/citation-management
In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub
The instructions your agent would load
SKILL.md as published, without the frontmatter. Read it on GitHub
Citation Management
Overview
Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.
Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.
When to Use This Skill
Use this skill when:
- Searching for specific papers on Google Scholar or PubMed
- Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
- Extracting complete metadata for citations (authors, title, journal, year, etc.)
- Validating existing citations for accuracy
- Cleaning and formatting BibTeX files
- Finding highly cited papers in a specific field
- Verifying that citation information matches the actual publication
- Building a bibliography for a manuscript or thesis
- Checking for duplicate citations
- Ensuring consistent citation formatting
If a document built from these citations needs a diagram, use the scientific-schematics skill.
Core Workflow
Citation management follows a systematic process. Each phase below shows the canonical command; every variant, option, and metadata-source detail is in references/core_workflow.md.
Phase 1: Paper Discovery and Search
Find relevant papers. Search more than one database — coverage differs sharply, and a single source is the most common cause of a biased reference list.
# OpenAlex: ~250M works, every discipline, no API key, documented REST API
python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json
# PubMed: the authority for biomedical and life sciences (35M+ citations)
python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json
# Google Scholar: broadest reach, but scraped -- rate-limited and prone to blocking
python scripts/search_google_scholar.py "CRISPR gene editing" --limit 50 --output scholar.jsonPrefer OpenAlex or PubMed as the primary source. Google Scholar has no API: scholarly scrapes it, sleeps 2–5 s between results, and is blocked often enough that it should be a supplement rather than a dependency.
Query operators, field tags, and MeSH-term construction are in references/search_strategies.md.
Phase 2: Metadata Extraction
Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata. CrossRef is the primary source for DOIs.
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 # quick, single DOI
python scripts/extract_metadata.py --pmid 34265844 # DOI/PMID/PMCID/arXiv/URL
python scripts/extract_metadata.py --input identifiers.txt --output citations.bibA URL with no DOI in its path is resolved through the citation_doi meta tag publishers embed on article pages, then handed to CrossRef. Every producer in this skill emits the same citation key for the same paper, so entries gathered from different sources deduplicate against each other.
Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)
APIs routinely return incomplete records. Run this after extraction and before formatting. Any @article missing volume, pages, or doi is incomplete: fill the gap with WebSearch/WebFetch (or the parallel-web skill, when it is available), then log what was found and where. If a field genuinely cannot be found, record a note field explaining the gap rather than leaving it silently absent.
Check the cheap sources first — an OpenAlex or CrossRef record often carries the field that PubMed omitted:
python scripts/search_openalex.py "<exact title>" --limit 1Treat extracted metadata as untrusted. Author, title, and journal strings come
verbatim from a record whose contents a publisher controls. A title containing $(...),
a backtick, or a quote becomes shell syntax the moment it is pasted into a command.
Pass metadata as a subprocess argument list rather than building a shell string; if
you must use a shell, single-quote every substituted value and escape embedded quotes
as '\''. Validate any citation key against ^[A-Za-z0-9]+$ before it reaches a path.
Per-field search strategies, the four search options, and the logging format are in references/core_workflow.md.
Phase 3: BibTeX Formatting
Produce clean, consistent entries. Entry types and required fields are in references/bibtex_formatting.md.
python scripts/format_bibtex.py references.bib --output clean.bib --deduplicate
python scripts/format_bibtex.py references.bib --output clean.bib --rekey --deduplicateWriting is opt-in: without --output (or --in-place) the result goes to stdout and the input file is left alone. Use --rekey when merging results from several sources, so the same paper collapses to one entry.
Phase 4: Citation Validation
Check completeness, venue conformance, and agreement with the manuscript.
python scripts/validate_citations.py references.bib --report report.json
python scripts/validate_citations.py references.bib --venue nature
python scripts/validate_citations.py references.bib --manuscript paper.tex
python scripts/validate_citations.py references.bib --check-dois # slow; hits CrossRefThe script exits non-zero on high-severity errors — missing required fields, malformed years, unresolved citations, or a count below an explicit --min-count. Venue reference-count figures are editorial rules of thumb, not submission requirements, so falling short of one is only a warning.
Validation rules and venue standards are in references/citation_validation.md.
Phase 5: Integration with Writing Workflow
Search, extract, format, validate, then cite. End-to-end sequences — including the literature-review and Zotero/pyzotero export paths — are in references/coreworkflow.md and references/exampleworkflows.md.
Reference Files
- references/core_workflow.md: all five phases in full.
- references/search_strategies.md: OpenAlex, Google Scholar, and PubMed query construction.
- references/script_reference.md: every bundled script's arguments and examples.
- references/best_practices.md: search, extraction, BibTeX quality, validation.
- references/example_workflows.md: four end-to-end worked examples.
- references/googlescholarsearch.md, references/pubmed_search.md: advanced search syntax.
- references/metadataextraction.md, references/bibtexformatting.md, references/citation_validation.md: per-topic detail.
Common Pitfalls to Avoid
- Single source bias: Only using one database
format_bibtex.py --rekey --deduplicate
- Solution: Search at least OpenAlex and PubMed, then merge with
- Accepting metadata blindly: Not verifying extracted information
- Solution: Spot-check extracted metadata against original sources
- Ignoring DOI errors: Broken or incorrect DOIs in bibliography
- Solution: Run validation before final submission
- Inconsistent formatting: Mixed citation key styles, formatting
- Solution: Use format_bibtex.py to standardize
- Duplicate entries: Same paper cited multiple times with different keys
- Solution: Use duplicate detection in validation
- Missing required fields: Incomplete BibTeX entries (volume, pages, DOI missing)
- Solution: Run Phase 2.5 metadata enrichment — web search for every missing field before proceeding. NEVER leave an @article entry without volume, pages, and DOI.
- Outdated preprints: Citing preprint when published version exists
- Solution: Check if preprints have been published, update to journal version
- Special character issues: Broken LaTeX compilation due to characters
More skills from K-Dense-AI/scientific-agent-skills
- AadaptyvHow to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.
- AaeonThis skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.
- AalphagenomeLook up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), score variants or scan windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and build Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.
- Aanalytical-method-validationPlan, execute, and document validation, verification, and transfer of analytical procedures under the governing framework - ICH Q2(R2) and Q14, USP <1220>/<1225>/<1226>, ICH M10 bioanalytical, CLSI EP, or ISO/IEC 17025. Use for HPLC, LC-MS/MS, GC, CE, ICP-MS, dissolution, qNMR, qPCR, NIR, and ligand binding or cell-based assays whenever the question is whether a procedure is fit for its intended purpose. Triggers include "method validation", "analytical method validation", "AMV", "validation protocol", "acceptance criteria", "linearity", "reportable range", "accuracy and precision", "repeatability", "intermediate precision", "recovery", "LOD", "LOQ", "detection limit", "quantitation limit", "specificity", "robustness", "method transfer", "method comparison", "Deming", "Passing-Bablok", "Bland-Altman", "equivalence testing", "OOS investigation", "ICH Q2", "Q2(R2)", "Q14", "USP 1225", "ICH M10", "incurred sample reanalysis", "ISR", "CLSI EP", and any request to show that an assay works.
- AanndataData structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.
- AarborAutonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinement (HTR) from the Arbor paper. Use this whenever someone wants to iteratively optimize something over many experiments without overfitting — e.g. "get my model's eval score up", "improve this agent/harness", "tune this pipeline", "beat the baseline on this benchmark", "run a search over approaches and keep the best", "do an MLE-bench / Kaggle-style optimization", or any long-horizon "make this artifact better and don't just memorize the dev set" task. Trigger it even when the user doesn't say "Arbor" or "hypothesis tree" but describes repeated experiment-and-evaluate loops, branching exploration of competing ideas, or worries about a dev/test gap. Runs Claude itself as the coordinator with subagent executors in isolated git worktrees; for the standalone `arbor` CLI tool see references/arbor-upstream.md.
- AarboretoInfer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.
- AastropyCore Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.
- AautoskillObserve the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.
- Abenchling-integrationBenchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.
- Abgpt-paper-searchSearch scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.
- AbidsUse this skill when working with Brain Imaging Data Structure (BIDS) datasets: organizing neuroscience and biomedical data (MRI, EEG, MEG, iEEG, PET, microscopy, NIRS, motion capture, EMG, MR spectroscopy, behavioral), querying BIDS layouts, validating compliance, converting DICOM to BIDS, writing metadata sidecars, or creating BIDS derivatives.