Mmcp.market

heartmula skill

by kevinnft·kevinnft/ai-agent-skills·14 stars·MIT

HeartMuLa: Suno-like song generation from lyrics + tags.

A100/100content scan

Is the heartmula skill safe?

Clean: nothing in its files matched our rules. We read 1 file in the folder on 2026-09-28.

No findings.

Install the heartmula skill

A skill is a folder. Copy it into your agent's skills folder and the agent loads it when the task matches its description.

git clone --depth 1 https://github.com/kevinnft/ai-agent-skills.git /tmp/ai-agent-skills
mkdir -p ~/.claude/skills
cp -r /tmp/ai-agent-skills/skills/media/heartmula ~/.claude/skills/heartmula
available in every project

In the Claude apps, zip the folder and upload it from the Skills settings. The folder on GitHub

The instructions your agent would load

SKILL.md as published, without the frontmatter. Read it on GitHub

HeartMuLa - Open-Source Music Generation

Overview

HeartMuLa is a family of open-source music foundation models (Apache-2.0) that generates music conditioned on lyrics and tags, with multilingual support. Generates full songs from lyrics + tags. Comparable to Suno for open-source. Includes:

  • HeartMuLa - Music language model (3B/7B) for generation from lyrics + tags
  • HeartCodec - 12.5Hz music codec for high-fidelity audio reconstruction
  • HeartTranscriptor - Whisper-based lyrics transcription
  • HeartCLAP - Audio-text alignment model

When to Use

  • User wants to generate music/songs from text descriptions
  • User wants an open-source Suno alternative
  • User wants local/offline music generation
  • User asks about HeartMuLa, heartlib, or AI music generation

Hardware Requirements

  • Minimum: 8GB VRAM with --lazy_load true (loads/unloads models sequentially)
  • Recommended: 16GB+ VRAM for comfortable single-GPU usage
  • Multi-GPU: Use --muladevice cuda:0 --codecdevice cuda:1 to split across GPUs
  • 3B model with lazy_load peaks at ~6.2GB VRAM

Installation Steps

1. Clone Repository

cd ~/  # or desired directory
git clone https://github.com/HeartMuLa/heartlib.git
cd heartlib

2. Create Virtual Environment (Python 3.10 required)

uv venv --python 3.10 .venv
. .venv/bin/activate
uv pip install -e .

3. Fix Dependency Compatibility Issues

IMPORTANT: As of Feb 2026, the pinned dependencies have conflicts with newer packages. Apply these fixes:

# Upgrade datasets (old version incompatible with current pyarrow)
uv pip install --upgrade datasets

# Upgrade transformers (needed for huggingface-hub 1.x compatibility)
uv pip install --upgrade transformers

4. Patch Source Code (Required for transformers 5.x)

Patch 1 - RoPE cache fix in src/heartlib/heartmula/modeling_heartmula.py:

In the setupcaches method of the HeartMuLa class, add RoPE reinitialization after the resetcaches try/except block and before the with device: block:

# Re-initialize RoPE caches that were skipped during meta-device loading
from torchtune.models.llama3_1._position_embeddings import Llama3ScaledRoPE
for module in self.modules():
    if isinstance(module, Llama3ScaledRoPE) and not module.is_cache_built:
        module.rope_init()
        module.to(device)

Why: frompretrained creates model on meta device first; Llama3ScaledRoPE.ropeinit() skips cache building on meta tensors, then never rebuilds after weights are loaded to real device.

Patch 2 - HeartCodec loading fix in src/heartlib/pipelines/music_generation.py:

Add ignoremismatchedsizes=True to ALL HeartCodec.frompretrained() calls (there are 2: the eager load in init__ and the lazy load in the codec property).

Why: VQ codebook initted buffers have shape [1] in checkpoint vs [] in model. Same data, just scalar vs 0-d tensor. Safe to ignore.

5. Download Model Checkpoints

cd heartlib  # project root
hf download --local-dir './ckpt' 'HeartMuLa/HeartMuLaGen'
hf download --local-dir './ckpt/HeartMuLa-oss-3B' 'HeartMuLa/HeartMuLa-oss-3B-happy-new-year'
hf download --local-dir './ckpt/HeartCodec-oss' 'HeartMuLa/HeartCodec-oss-20260123'

All 3 can be downloaded in parallel. Total size is several GB.

GPU / CUDA

HeartMuLa uses CUDA by default (--muladevice cuda --codecdevice cuda). No extra setup needed if the user has an NVIDIA GPU with PyTorch CUDA support installed.

  • The installed torch==2.4.1 includes CUDA 12.1 support out of the box
  • torchtune may report version 0.4.0+cpu — this is just package metadata, it still uses CUDA via PyTorch
  • To verify GPU is being used, look for "CUDA memory" lines in the output (e.g. "CUDA memory before unloading: 6.20 GB")
  • No GPU? You can run on CPU with --muladevice cpu --codecdevice cpu, but expect generation to be extremely slow (potentially 30-60+ minutes for a single song vs ~4 minutes on GPU). CPU mode also requires significant RAM (~12GB+ free). If the user has no NVIDIA GPU, recommend using a cloud GPU service (Google Colab free tier with T4, Lambda Labs, etc.) or the online demo at https://heartmula.github.io/ instead.

Usage

Basic Generation

cd heartlib
. .venv/bin/activate
python ./examples/run_music_generation.py \
  --model_path=./ckpt \
  --version="3B" \
  --lyrics="./assets/lyrics.txt" \
  --tags="./assets/tags.txt" \
  --save_path="./assets/output.mp3" \
  --lazy_load true

Input Formatting

Tags (comma-separated, no spaces):

piano,happy,wedding,synthesizer,romantic

or

rock,energetic,guitar,drums,male-vocal

Lyrics (use bracketed structural tags):

[Intro]

[Verse]
Your lyrics here...

[Chorus]
Chorus lyrics...

[Bridge]
Bridge lyrics...

[Outro]

Key Parameters

Performance

  • RTF (Real-Time Factor) ≈ 1.0 — a 4-minute song takes ~4 minutes to generate
  • Output: MP3, 48kHz stereo, 128kbps

Pitfalls

  1. Do NOT use bf16 for HeartCodec — degrades audio quality. Use fp32 (default).
  2. Tags may be ignored — known issue (#90). Lyrics tend to dominate; experiment with tag ordering.
  3. Triton not available on macOS — Linux/CUDA only for GPU acceleration.
  4. RTX 5080 incompatibility reported in upstream issues.
  5. The dependency pin conflicts require the manual upgrades and patches described above.

Links

  • Repo: https://github.com/HeartMuLa/heartlib
  • Models: https://huggingface.co/HeartMuLa
  • Paper: https://arxiv.org/abs/2601.10547
  • License: Apache-2.0

More skills from kevinnft/ai-agent-skills

  • Aaddyosmani-tddDrives development with tests. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.
  • AairtableAirtable REST API via curl. Records CRUD, filters, upserts.
  • Aapi-and-interface-designGuides stable API and interface design. Use when designing APIs, module boundaries, or any public interface. Use when creating REST or GraphQL endpoints, defining type contracts between modules, or establishing boundaries between frontend and backend.
  • Aapi-monitoring-botsBuild monitoring bots that poll APIs and send notifications on state changes (new listings, price alerts, status updates)
  • Aapple-notesManage Apple Notes via memo CLI: create, search, edit.
  • Aapple-remindersApple Reminders via remindctl: add, list, complete.
  • Aarchitecture-diagramDark-themed SVG architecture/cloud/infra diagrams as HTML.
  • AarxivSearch arXiv papers by keyword, author, category, or ID.
  • Aascii-artASCII art: pyfiglet, cowsay, boxes, image-to-ascii.
  • Aascii-videoASCII video: convert video/audio to colored ASCII MP4/GIF.
  • AaudiocraftAudioCraft: MusicGen text-to-music, AudioGen text-to-sound.
  • CaxolotlAxolotl: YAML LLM fine-tuning (LoRA, DPO, GRPO).

All agent skills → · MCP servers