Importing Model Weights
Import safetensors weights from Hugging Face and build custom local models with a Modelfile and ollama create.
TL;DR
- Download
.safetensorsweights from Hugging Face model hubs first. - Point
FROMat the directory holding the weights. - Build and quantize a local model with
ollama create.
Get Weights From Hugging Face
git lfsInstall Git LFS to fetch large weight files.
git lfs installClone A RepoClone a model repo with safetensors and config.
git clone https://huggingface.co/org/modelRequired FilesA valid folder has safetensors plus config.json.
ls model/
# *.safetensors config.json tokenizer.jsonModelfile FROM Directory
FROM .Point at the current folder holding the weights.
FROM .FROM ./dirOr point at a subdirectory of safetensors weights.
FROM ./my-modelAdd ParametersOptionally set default parameters or a template.
FROM .
PARAMETER temperature 0.7Create And Quantize
ollama createBuild the model from the Modelfile in this folder.
ollama create my-model--quantizeQuantize an fp16 safetensors model during create.
ollama create my-model --quantize q4_K_MRun ItStart the newly built local model.
ollama run my-modelImport GGUF Directly
FROM file.ggufPoint FROM at a single prebuilt GGUF file.
FROM ./model.q4_k_m.ggufSplit FilesUse a wildcard for multi-part GGUF downloads.
FROM ./model-*.ggufNo Re-QuantizeOllama imports GGUF as-is without changing precision.
# Pre-quantize with llama-quantize firstTips
- Keep the
Modelfilein the same folder as the weights, then useFROM .so the path never breaks when you move the project. - Add
--quantize q4_K_Mtoollama createto shrink a full-precision safetensors model as you import it, saving disk and memory.
Warnings
FROMmust point to the safetensors directory, not a single.safetensorsfile; the folder also needsconfig.jsonbeside the weights.- Ollama does not re-quantize
GGUFfiles on import; pre-quantize those with a tool likellama-quantizebefore pointingFROMat them.
In Practice
Download safetensors weights, write a one-line Modelfile, then build and quantize a custom local model.
- Git LFS pulls the large safetensors files the model needs.
FROM .points at the folder so the weights load from beside the Modelfile.--quantize q4_K_Mshrinks the fp16 weights as Ollama imports them.- The finished model runs like any other local model.
# 1. Download the weights (safetensors + config)
git lfs install
git clone https://huggingface.co/org/my-model
# 2. Create a Modelfile in that folder
cd my-model
echo 'FROM .' > Modelfile
# 3. Build and quantize in one step
ollama create my-model --quantize q4_K_M
# 4. Run it
ollama run my-modelFAQ
First clone the model repo so you have its safetensors weights and config.json. Create a Modelfile containing FROM . in that folder, then run ollama create my-model. Ollama converts the weights into its own format.
For safetensors, FROM points to the directory containing the weights and config.json, not a single file. For a prebuilt GGUF, FROM points to the .gguf file itself.
Yes, for fp16 or fp32 safetensors models. Add --quantize q4_K_M to ollama create and Ollama quantizes during import. It cannot re-quantize an existing GGUF file, though.
Write a one-line Modelfile with FROM ./model.gguf, then run ollama create. For a model split into parts, use a wildcard like FROM ./model-*.gguf instead.