Skill

Estimate Model Memory Requirements Without Downloading

A CLI skill for estimating Hugging Face model inference memory - weights and KV cache, no download required.

Works with huggingface

84
Spark score
out of 100
Updated 23 days ago
Version 1.0.0

Add to Favorites

Why it matters

Help developers and ML engineers determine the exact VRAM and memory requirements for running AI models from Hugging Face Hub before deployment, enabling informed decisions about hardware provisioning and instance selection without downloading gigabytes of model weights.

Outcomes

What it gets done

01

Calculate inference memory for Safetensors and GGUF models using HTTP Range requests

02

Estimate KV cache memory requirements for LLMs and VLMs with configurable context windows

03

Check if a specific model fits on available GPU or instance memory

04

Compare memory needs across different quantization levels and precisions

Install

Add it to your toolbox

Run in your project directory:

curl -fsSL https://spark.entire.vc/get/ag-hf-mem | bash

Overview

Hf Mem

This skill uses hf-mem to estimate Hugging Face model inference memory - Safetensors or GGUF weights, plus optional KV cache for LLMs/VLMs - via HTTP Range requests without downloading weights. Use it when a user asks how much VRAM or memory a Hugging Face model needs, or whether it fits their GPU or instance.

What it does

A skill for hf_mem, a CLI tool that estimates the memory required for inference - model weights plus an optional KV cache - for Safetensors and GGUF models on the Hugging Face Hub, using HTTP Range requests so no weights are downloaded or loaded locally. uvx hf-mem --model-id <model-id> --json-output checks that the repo contains Safetensors (model.safetensors, a sharded model.safetensors.index.json, or model_index.json for Diffusers) or GGUF weights. For GGUF repos with multiple precisions/quantizations, estimation is per-file, so --gguf-file <file-or-path> targets the specific precision actually intended for inference rather than estimating every variant. An --experimental flag adds KV cache memory estimation, applicable to LLMs (...ForCausalLM), VLMs (...ForConditionalGeneration), and GGUF models - the context window is read from the model default or overridden with --max-model-len (vLLM-style), and KV cache precision defaults to the model's own precision unless set via --kv-cache-dtype, which for Safetensors accepts auto/bfloat16/fp8/fp8_ds_mla/fp8_e4m3/fp8_e5m2/fp8_inc and for GGUF accepts a much longer list of quantization types (F32, F16, Q4_0 through Q8_K, the IQ family, BF16, TQ1_0/TQ2_0, MXFP4, and more). Documented examples cover Transformers models with Safetensors weights (MiniMaxAI/MiniMax-M2), Diffusers models (Qwen/Qwen-Image), Sentence Transformers (google/embeddinggemma-300m), an LLM with KV cache estimation via --experimental (mistralai/Mistral-7B-v0.1), and a GGUF LLM/VLM targeting a specific quantized file (unsloth/Qwen3.5-397B-A17B-GGUF --gguf-file Q4_K_M --experimental).

When to use - and when NOT to

Use it when a user asks how much VRAM or memory a model needs to run, whether a model fits on their GPU or a given instance, or references a Hugging Face model ID/URL and asks about inference requirements. Requires uv installed (for uvx) and an HF_TOKEN environment variable or --hf-token flag for gated or private models only.

Inputs and outputs

Input is a Hugging Face model ID (and for GGUF, the specific file/precision to target) plus optional flags for KV cache estimation, context length, batch size, and cache dtype. Output is a JSON report of estimated memory requirements for model weights and, with --experimental, KV cache.

Integrations

uvx hf-mem --model-id <model-id> --json-output

Reads model metadata directly from the Hugging Face Hub via HTTP Range requests - no local download, no GPU or inference framework required to run the estimation itself.

Who it's for

Developers and ML practitioners deciding whether a Hugging Face model fits their available VRAM before downloading or deploying it - checking Safetensors or GGUF weight memory alone, or including KV cache requirements for LLM/VLM inference at a given context length and precision.

Source README

hf_mem estimates the required memory for inference, including model weights and an optional KV cache, for Safetensors and GGUF for models on the Hugging Face Hub using HTTP Range requests i.e., without downloading or loading any weights locally.

When to use?

  • User asks how much VRAM or memory a model needs to run
  • User wants to know if a model fits on their GPU or a given instance
  • User references a Hugging Face model ID or URL and asks about inference requirements

What are the requirements?

  • uv installed (for uvx)
  • HF_TOKEN env var or --hf-token flag (for gated or private models only)

How to run?

Run with --model-id pointing to the Hugging Face Hub repository which will check that it either contains Safetensors (via model.safetensors, model.safetensors.index.json if sharded, or model_index.json for Diffusers) or GGUF model weights within.

uvx hf-mem --model-id <model-id> --json-output

If the repository contains GGUF model weights in multiple precisions / quantizations, the estimations will be on a per-file basis, whereas for inference you won't load all of those but rather only a single precision. This being said, for GGUF you might as well need to provide --gguf-file to target the specific file (or path if sharded) you want to run.

uvx hf-mem --model-id <model-id> --gguf-file <file-or-path> --json-output

Additionally, hf-mem comes with an --experimental flag that will also calculate the KV cache memory requirements too, useful for large-language models, meaning it applies to LLMs (...ForCausalLM), VLMs (...ForConditionalGeneration), and GGUF models.

As per the context window, it will be read from the default or overridden with --max-model-len a la vLLM. And, same goes for the KV cache precision, which will default to the model precision unless manually set via --kv-cache-dtype a la vLLM too.

For Safetensors use as:

uvx hf-mem --model-id <model-id> --experimental [--max-model-len N] [--batch-size N] [--kv-cache-dtype auto|bfloat16|fp8|fp8_ds_mla|fp8_e4m3|fp8_e5m2|fp8_inc] --json-output

And, for GGUF use as:

uvx hf-mem --model-id <model-id> --gguf-file <file-or-path> --experimental [--max-model-len N] [--batch-size N] [--kv-cache-dtype auto|F32|F16|Q4_0|Q4_1|Q5_0|Q5_1|Q8_0|Q8_1|Q2_K|Q3_K|Q4_K|Q5_K|Q6_K|Q8_K|IQ2_XXS|IQ2_XS|IQ3_XXS|IQ1_S|IQ4_NL|IQ3_S|IQ2_S|IQ4_XS|I8|I16|I32|I64|F64|IQ1_M|BF16|TQ1_0|TQ2_0|MXFP4] --json-output

Examples

For Transformers with Safetensors weights:

uvx hf-mem --model-id MiniMaxAI/MiniMax-M2 --json-output

For Diffusers with Safetensors weights:

uvx hf-mem --model-id Qwen/Qwen-Image --json-output

For Sentence Transformers with Safetensors weights:

uvx hf-mem --model-id google/embeddinggemma-300m --json-output

With --experimental to include the KV cache estimation for LLMs and VLMs:

uvx hf-mem --model-id mistralai/Mistral-7B-v0.1 --experimental --json-output

And, for LLMs or VLMs with GGUF weights:

uvx hf-mem --model-id unsloth/Qwen3.5-397B-A17B-GGUF --gguf-file Q4_K_M --experimental --json-output

Limitations

  • Use this skill only when the task clearly matches its upstream product or API scope.
  • Verify commands, API behavior, pricing, quotas, credentials, and deployment effects against current official documentation before making changes.
  • Do not treat generated examples as a substitute for environment-specific tests, security review, or user approval for destructive or costly actions.

FAQ

Common questions

Discussion

Questions & comments ยท 0

Sign In Sign in to leave a comment.