Importing Model Weights

Import safetensors weights from Hugging Face and build custom local models with a Modelfile and ollama create.

TL;DR

  1. Download .safetensors weights from Hugging Face model hubs first.
  2. Point FROM at the directory holding the weights.
  3. Build and quantize a local model with ollama create.

Get Weights From Hugging Face

    git lfs

    Install Git LFS to fetch large weight files.

    git lfs install
    Clone A Repo

    Clone a model repo with safetensors and config.

    git clone https://huggingface.co/org/model
    Required Files

    A valid folder has safetensors plus config.json.

    ls model/
    # *.safetensors  config.json  tokenizer.json

Modelfile FROM Directory

    FROM .

    Point at the current folder holding the weights.

    FROM .
    FROM ./dir

    Or point at a subdirectory of safetensors weights.

    FROM ./my-model
    Add Parameters

    Optionally set default parameters or a template.

    FROM .
    PARAMETER temperature 0.7

Create And Quantize

    ollama create

    Build the model from the Modelfile in this folder.

    ollama create my-model
    --quantize

    Quantize an fp16 safetensors model during create.

    ollama create my-model --quantize q4_K_M
    Run It

    Start the newly built local model.

    ollama run my-model

Import GGUF Directly

    FROM file.gguf

    Point FROM at a single prebuilt GGUF file.

    FROM ./model.q4_k_m.gguf
    Split Files

    Use a wildcard for multi-part GGUF downloads.

    FROM ./model-*.gguf
    No Re-Quantize

    Ollama imports GGUF as-is without changing precision.

    # Pre-quantize with llama-quantize first

Tips

  1. Keep the Modelfile in the same folder as the weights, then use FROM . so the path never breaks when you move the project.
  2. Add --quantize q4_K_M to ollama create to shrink a full-precision safetensors model as you import it, saving disk and memory.

Warnings

  1. FROM must point to the safetensors directory, not a single .safetensors file; the folder also needs config.json beside the weights.
  2. Ollama does not re-quantize GGUF files on import; pre-quantize those with a tool like llama-quantize before pointing FROM at them.

In Practice

FAQ