Hugging Face Cache Trap: Hidden Folder Eating Server Disk
Your AI app can die overnight when ~/.cache/huggingface fills the root disk. Learn the Australian server signs, DIY cleanup commands, and when Fixwebnode should step in remotely.
If you run transformers, diffusers, or any Hugging Face model stack on a Linux VPS or bare-metal box in Australia, one invisible folder can exhaust root storage and take the whole app down without a clear error. This guide walks sysadmins and small-business owners through finding the cache trap, reclaiming space safely, and hardening paths so it does not happen again.
Fixwebnode provides direct remote Linux and server support for this exact failure mode—disk full, model downloads looping, and broken inference services—not a freelance marketplace. Start with the practical steps below, or book remote help via Ubuntu Linux server support when the box is already in production pain.
Why the Hugging Face cache trap matters on Australian servers
Hugging Face libraries default to caching model weights, tokenizers, and datasets under the user home directory—commonly ~/.cache/huggingface or $HF_HOME. On a typical cloud VPS the root volume is modest. A few large LLM or diffusion checkpoints quietly push / past 95% full. Cron jobs fail, Docker cannot create layers, PostgreSQL refuses writes, and your AI API returns 500s or hangs on load.
The folder is easy to miss in casual df checks because people look at project directories, not hidden cache trees. Australian teams running staging and production on the same small instance feel this first: overnight model pulls after a deploy fill the disk before morning stand-up.
What fills the Hugging Face cache and crashes AI apps in Australia?
The Hugging Face cache trap is a default download directory (often under the home folder) that stores model weights and datasets with no automatic size cap. On Australian VPS and dedicated servers it commonly fills the root disk until inference, deploys, and even SSH sessions degrade or fail. Fix it by measuring disk use, relocating or pruning the cache, and setting explicit environment limits—or book Fixwebnode for remote recovery if the host is already read-only.
| Symptom | Quick check | When to call Fixwebnode |
|---|---|---|
| Root disk 95–100% full | du -sh ~/.cache/huggingface | Host is read-only or SSH unstable |
| Model load / OOM after “download complete” | Inspect hub blobs and incomplete files | Production API down across restarts |
| Cache returns after every deploy | Check HF_HOME / Docker volume mounts | Need permanent path + systemd hardening |
Common issues caused by the hidden Hugging Face cache
These problems share the same root cause—unbounded cache growth—but show up differently in logs and monitoring.
- Root filesystem silently full.
df -h /shows 100% used; new writes fail with “No space left on device.” Apps crash; package updates abort. - Repeated multi‑GB downloads on every restart. Incomplete or purged blobs force Hugging Face Hub to re-fetch weights, saturating bandwidth and CPU while the API stays cold.
- Permission and multi-user cache collisions. systemd services, Docker containers, and interactive users each write different cache trees (or fight over one), leaving root-owned blobs the app user cannot read.
- Docker / Compose layer and volume bloat. Images bake cache into layers, or anonymous volumes grow under
/var/lib/dockeruntil the host disk dies even after you cleaned~/.cache.
Issue 1 — Root disk full from ~/.cache/huggingface
Symptoms: sudden “No space left on device,” failed apt operations, databases dropping writes, AI workers exiting. df -h points at /; project folders look small.
Step 1 — Confirm disk pressure and locate the cache
df -h /
df -i /
du -xh /home /root /var 2>/dev/null | sort -h | tail -n 30
du -sh /home/*/.cache/huggingface /root/.cache/huggingface 2>/dev/null
ls -la ~/.cache/huggingface 2>/dev/null
env | grep -E 'HF_|TRANSFORMERS_'
Large hub, transformers, or datasets directories under the Hugging Face cache confirm the trap. Inodes matter too: millions of tiny blob files can exhaust inodes before gigabytes look dramatic.
Step 2 — Measure the worst offenders inside the cache
du -h --max-depth=2 ~/.cache/huggingface 2>/dev/null | sort -h
find ~/.cache/huggingface -type f -size +500M -printf '%s\t%p\n' 2>/dev/null | sort -n
Step 3 — Safe DIY reclaim (stop writers first)
sudo systemctl stop your-ai-app.service 2>/dev/null || true
# If using Docker Compose for the model service:
cd /opt/your-app && sudo docker compose stop
# Delete only stale Hub blobs / incomplete downloads after review
rm -rf ~/.cache/huggingface/hub/.locks
find ~/.cache/huggingface -type f -name '*.incomplete' -delete
# Optional full cache wipe when models can be re-downloaded:
# rm -rf ~/.cache/huggingface/*
df -h /
Prefer deleting incomplete files and locks before a full wipe. After space returns, restart the service and watch first-boot download size.
Step 4 — Verify
df -h /
du -sh ~/.cache/huggingface
journalctl -u your-ai-app.service -n 50 --no-pager
When to call Fixwebnode: if the root filesystem is already read-only, you cannot log in cleanly, or critical databases share the same volume and you cannot risk aggressive deletes. Remote recovery can free space, move cache off root, and stabilise services without guesswork.
Issue 2 — Models re-download every deploy and thrash the disk
Symptoms: long cold starts, Hub rate limits, bandwidth spikes, cache growing again within hours even after cleanup. Logs show repeated “Downloading …” for the same revision.
Step 1 — Pin a durable cache location outside a tiny root volume
sudo mkdir -p /var/cache/huggingface
sudo chown -R deploy:deploy /var/cache/huggingface
# Example for a service user named deploy
grep -E 'HF_HOME|HUGGINGFACE_HUB_CACHE' /etc/environment /etc/systemd/system/*.service 2>/dev/null
Step 2 — Set environment variables for shell and systemd
# Interactive / deploy user
echo 'export HF_HOME=/var/cache/huggingface' >> ~/.bashrc
echo 'export HUGGINGFACE_HUB_CACHE=/var/cache/huggingface/hub' >> ~/.bashrc
echo 'export TRANSFORMERS_CACHE=/var/cache/huggingface/transformers' >> ~/.bashrc
source ~/.bashrc
# systemd drop-in example
sudo mkdir -p /etc/systemd/system/your-ai-app.service.d
sudo tee /etc/systemd/system/your-ai-app.service.d/hf-cache.conf <<'EOF'
[Service]
Environment=HF_HOME=/var/cache/huggingface
Environment=HUGGINGFACE_HUB_CACHE=/var/cache/huggingface/hub
Environment=TRANSFORMERS_CACHE=/var/cache/huggingface/transformers
Environment=HF_HUB_DISABLE_XET=1
EOF
sudo systemctl daemon-reload
sudo systemctl restart your-ai-app.service
Step 3 — Optional: move existing cache without re-downloading
sudo systemctl stop your-ai-app.service
rsync -aH --info=progress2 ~/.cache/huggingface/ /var/cache/huggingface/
mv ~/.cache/huggingface ~/.cache/huggingface.bak.$(date +%F)
ln -s /var/cache/huggingface ~/.cache/huggingface
sudo systemctl start your-ai-app.service
du -sh /var/cache/huggingface
Step 4 — Verify offline / local load behaviour
# After models are present, prefer local files in app config where supported
# e.g. local_files_only=True in transformers, or HF_HUB_OFFLINE=1 for a test run
HF_HUB_OFFLINE=1 systemctl restart your-ai-app.service
journalctl -u your-ai-app.service -n 100 --no-pager | tail
If the service still hits the network, a code path or container is ignoring HF_HOME. Trace the unit file, Compose environment: block, and any entrypoint scripts.
When to call Fixwebnode: when multiple services (API, worker, notebook, CI runner) each maintain separate caches, or deploys are automated and keep resetting env vars. A specialist can unify paths and bake durable settings into systemd and Compose.
Issue 3 — Permission clashes and root-owned cache blobs
Symptoms: app user logs “Permission denied” under .cache/huggingface/hub; manual sudo python tests work but the service fails; mixed UIDs inside the cache tree.
Step 1 — Audit ownership
sudo find /var/cache/huggingface ~/.cache/huggingface -xdev -printf '%u:%g %p\n' 2>/dev/null | sort | uniq -c | sort -nr | head
ps aux | grep -E 'python|uvicorn|gunicorn' | grep -v grep
systemctl show your-ai-app.service -p User -p Group -p Environment
Step 2 — Fix ownership and stop running Hub clients as root
sudo systemctl stop your-ai-app.service
sudo chown -R deploy:deploy /var/cache/huggingface
# Ensure the unit runs as the same non-root user
sudo systemctl edit your-ai-app.service
# Add under [Service]: User=deploy and Group=deploy if missing
sudo systemctl daemon-reload
sudo systemctl start your-ai-app.service
Step 3 — Verify read/write as the service user
sudo -u deploy touch /var/cache/huggingface/.write_test && sudo -u deploy rm /var/cache/huggingface/.write_test
sudo -u deploy ls /var/cache/huggingface/hub | head
When to call Fixwebnode: if containers run as root while bind-mounting a host cache used by a non-root host process, or SELinux/AppArmor denials appear. Wrong ownership fixes often bounce back until unit files and Compose users are aligned.
Issue 4 — Docker hides a second Hugging Face cache under /var/lib/docker
Symptoms: you cleaned ~/.cache but df barely moved; /var/lib/docker is enormous; containers re-download models on every recreate.
Step 1 — Diagnose Docker disk use
df -h / /var/lib/docker
sudo du -sh /var/lib/docker/* 2>/dev/null | sort -h
sudo docker system df
sudo docker system df -v | head -n 80
Step 2 — Mount a host cache into the container (Compose example)
# docker-compose.yml fragment (illustrative)
# services:
# api:
# environment:
# HF_HOME: /cache/huggingface
# volumes:
# - /var/cache/huggingface:/cache/huggingface
# user: "1000:1000"
Apply the same UID on host and container where possible. Recreate—not only restart—so environment and mounts attach:
cd /opt/your-app
sudo docker compose up -d --force-recreate api
sudo docker compose exec api printenv | grep HF_
sudo docker compose exec api du -sh /cache/huggingface
Step 3 — Prune unused build cache carefully
# Review first; this removes unused data, not running container volumes you still need
sudo docker builder prune -f
sudo docker system prune -f
df -h /
Do not run aggressive volume prune on production until you confirm named volumes that hold only disposable cache.
When to call Fixwebnode: when production Compose stacks share anonymous volumes, or disk pressure sits under the Docker data root on a single root partition and you need a safe migration layout.
Hardening checklist so the cache cannot fill root again
After emergency cleanup, lock in guardrails.
- Point
HF_HOMEat a dedicated disk or large partition (for example/var/cache/huggingfaceor a mounted data volume). - Add filesystem monitoring: alert at 80% and 90% on
/and the cache mount. - Document which models are required; delete unused Hub snapshots deliberately.
- Keep systemd/Compose env vars in version control so deploys do not revert to home-directory defaults.
- Avoid downloading models as root in ad-hoc SSH sessions.
# Simple daily size report (crontab -e as deploy or root mail target)
# 15 7 * * * df -h / /var/cache 2>/dev/null; du -sh /var/cache/huggingface 2>/dev/null
# Snapshot of top cache revisions for audit
du -h --max-depth=3 /var/cache/huggingface/hub 2>/dev/null | sort -h | tail -n 40
When DIY is enough vs when to book Fixwebnode
DIY is enough when you still have SSH, disk is not fully read-only, you can stop the AI service briefly, and a single user owns the cache. The commands above clear incomplete downloads, relocate HF_HOME, fix ownership, and bind-mount Docker caches.
Book Fixwebnode when the server is already degraded (read-only root, failed multi-service stack, database corruption risk on the same volume), when several environments fight over cache paths, or when you need a durable layout across systemd, Docker, and deploy pipelines without extended downtime. Support is direct remote specialist work for Australian businesses and self-hosted stacks—not a bid board.
Geography: remote Linux and server recovery is available across Australia; see all service areas for coverage context. For this cache and disk-full class of incidents, remote sessions are usually the fastest path.
Talk to Fixwebnode about Hugging Face disk recovery
If the invisible Hugging Face cache has already pushed your root disk to the edge—or you want the paths, systemd units, and Docker mounts set correctly before the next model drop—start a conversation with Fixwebnode. Bring your df -h output and service names; we will work the live host directly.
Book remote Ubuntu and Linux server support for this issue via the landing page: https://fixwebnode.com.au/ubuntu-linux-server-support-bug-fixing-adelaide-australia. Get the cache off the root volume, restore inference, and leave monitoring in place so the hidden folder cannot quietly crash the app again.