Estimate Model Memory Requirements Without Downloading
A CLI skill for estimating Hugging Face model inference memory - weights and KV cache, no download required.
Why it matters
Help developers and ML engineers determine the exact VRAM and memory requirements for running AI models from Hugging Face Hub before deployment, enabling informed decisions about hardware provisioning and instance selection without downloading gigabytes of model weights.
Outcomes
What it gets done
Calculate inference memory for Safetensors and GGUF models using HTTP Range requests
Estimate KV cache memory requirements for LLMs and VLMs with configurable context windows
Check if a specific model fits on available GPU or instance memory
Compare memory needs across different quantization levels and precisions
Install
Add it to your toolbox
Run in your project directory:
curl -fsSL https://spark.entire.vc/get/ag-hf-mem | bash Overview
Hf Mem
This skill uses hf-mem to estimate Hugging Face model inference memory - Safetensors or GGUF weights, plus optional KV cache for LLMs/VLMs - via HTTP Range requests without downloading weights. Use it when a user asks how much VRAM or memory a Hugging Face model needs, or whether it fits their GPU or instance.
What it does
A skill for hf_mem, a CLI tool that estimates the memory required for inference - model weights plus an optional KV cache - for Safetensors and GGUF models on the Hugging Face Hub, using HTTP Range requests so no weights are downloaded or loaded locally. uvx hf-mem --model-id <model-id> --json-output checks that the repo contains Safetensors (model.safetensors, a sharded model.safetensors.index.json, or model_index.json for Diffusers) or GGUF weights. For GGUF repos with multiple precisions/quantizations, estimation is per-file, so --gguf-file <file-or-path> targets the specific precision actually intended for inference rather than estimating every variant. An --experimental flag adds KV cache memory estimation, applicable to LLMs (...ForCausalLM), VLMs (...ForConditionalGeneration), and GGUF models - the context window is read from the model default or overridden with --max-model-len (vLLM-style), and KV cache precision defaults to the model's own precision unless set via --kv-cache-dtype, which for Safetensors accepts auto/bfloat16/fp8/fp8_ds_mla/fp8_e4m3/fp8_e5m2/fp8_inc and for GGUF accepts a much longer list of quantization types (F32, F16, Q4_0 through Q8_K, the IQ family, BF16, TQ1_0/TQ2_0, MXFP4, and more). Documented examples cover Transformers models with Safetensors weights (MiniMaxAI/MiniMax-M2), Diffusers models (Qwen/Qwen-Image), Sentence Transformers (google/embeddinggemma-300m), an LLM with KV cache estimation via --experimental (mistralai/Mistral-7B-v0.1), and a GGUF LLM/VLM targeting a specific quantized file (unsloth/Qwen3.5-397B-A17B-GGUF --gguf-file Q4_K_M --experimental).
When to use - and when NOT to
Use it when a user asks how much VRAM or memory a model needs to run, whether a model fits on their GPU or a given instance, or references a Hugging Face model ID/URL and asks about inference requirements. Requires uv installed (for uvx) and an HF_TOKEN environment variable or --hf-token flag for gated or private models only.
Inputs and outputs
Input is a Hugging Face model ID (and for GGUF, the specific file/precision to target) plus optional flags for KV cache estimation, context length, batch size, and cache dtype. Output is a JSON report of estimated memory requirements for model weights and, with --experimental, KV cache.
Integrations
uvx hf-mem --model-id <model-id> --json-output
Reads model metadata directly from the Hugging Face Hub via HTTP Range requests - no local download, no GPU or inference framework required to run the estimation itself.
Who it's for
Developers and ML practitioners deciding whether a Hugging Face model fits their available VRAM before downloading or deploying it - checking Safetensors or GGUF weight memory alone, or including KV cache requirements for LLM/VLM inference at a given context length and precision.
Source README
hf_mem estimates the required memory for inference, including model weights and an optional KV cache, for Safetensors and GGUF for models on the Hugging Face Hub using HTTP Range requests i.e., without downloading or loading any weights locally.
When to use?
- User asks how much VRAM or memory a model needs to run
- User wants to know if a model fits on their GPU or a given instance
- User references a Hugging Face model ID or URL and asks about inference requirements
What are the requirements?
uvinstalled (foruvx)HF_TOKENenv var or--hf-tokenflag (for gated or private models only)
How to run?
Run with --model-id pointing to the Hugging Face Hub repository which will check that it either contains Safetensors (via model.safetensors, model.safetensors.index.json if sharded, or model_index.json for Diffusers) or GGUF model weights within.
uvx hf-mem --model-id <model-id> --json-output
If the repository contains GGUF model weights in multiple precisions / quantizations, the estimations will be on a per-file basis, whereas for inference you won't load all of those but rather only a single precision. This being said, for GGUF you might as well need to provide --gguf-file to target the specific file (or path if sharded) you want to run.
uvx hf-mem --model-id <model-id> --gguf-file <file-or-path> --json-output
Additionally, hf-mem comes with an --experimental flag that will also calculate the KV cache memory requirements too, useful for large-language models, meaning it applies to LLMs (...ForCausalLM), VLMs (...ForConditionalGeneration), and GGUF models.
As per the context window, it will be read from the default or overridden with --max-model-len a la vLLM. And, same goes for the KV cache precision, which will default to the model precision unless manually set via --kv-cache-dtype a la vLLM too.
For Safetensors use as:
uvx hf-mem --model-id <model-id> --experimental [--max-model-len N] [--batch-size N] [--kv-cache-dtype auto|bfloat16|fp8|fp8_ds_mla|fp8_e4m3|fp8_e5m2|fp8_inc] --json-output
And, for GGUF use as:
uvx hf-mem --model-id <model-id> --gguf-file <file-or-path> --experimental [--max-model-len N] [--batch-size N] [--kv-cache-dtype auto|F32|F16|Q4_0|Q4_1|Q5_0|Q5_1|Q8_0|Q8_1|Q2_K|Q3_K|Q4_K|Q5_K|Q6_K|Q8_K|IQ2_XXS|IQ2_XS|IQ3_XXS|IQ1_S|IQ4_NL|IQ3_S|IQ2_S|IQ4_XS|I8|I16|I32|I64|F64|IQ1_M|BF16|TQ1_0|TQ2_0|MXFP4] --json-output
Examples
For Transformers with Safetensors weights:
uvx hf-mem --model-id MiniMaxAI/MiniMax-M2 --json-output
For Diffusers with Safetensors weights:
uvx hf-mem --model-id Qwen/Qwen-Image --json-output
For Sentence Transformers with Safetensors weights:
uvx hf-mem --model-id google/embeddinggemma-300m --json-output
With --experimental to include the KV cache estimation for LLMs and VLMs:
uvx hf-mem --model-id mistralai/Mistral-7B-v0.1 --experimental --json-output
And, for LLMs or VLMs with GGUF weights:
uvx hf-mem --model-id unsloth/Qwen3.5-397B-A17B-GGUF --gguf-file Q4_K_M --experimental --json-output
Limitations
- Use this skill only when the task clearly matches its upstream product or API scope.
- Verify commands, API behavior, pricing, quotas, credentials, and deployment effects against current official documentation before making changes.
- Do not treat generated examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.
FAQ
Common questions
Discussion
Questions & comments ยท 0
Sign In Sign in to leave a comment.