Instructions to use TheFinAI/FinLLaMA-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TheFinAI/FinLLaMA-instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="TheFinAI/FinLLaMA-instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("TheFinAI/FinLLaMA-instruct") model = AutoModelForCausalLM.from_pretrained("TheFinAI/FinLLaMA-instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use TheFinAI/FinLLaMA-instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "TheFinAI/FinLLaMA-instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/TheFinAI/FinLLaMA-instruct
- SGLang
How to use TheFinAI/FinLLaMA-instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "TheFinAI/FinLLaMA-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "TheFinAI/FinLLaMA-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "TheFinAI/FinLLaMA-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use TheFinAI/FinLLaMA-instruct with Docker Model Runner:
docker model run hf.co/TheFinAI/FinLLaMA-instruct
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="TheFinAI/FinLLaMA-instruct")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages)# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("TheFinAI/FinLLaMA-instruct")
model = AutoModelForCausalLM.from_pretrained("TheFinAI/FinLLaMA-instruct", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))FinLLaMA-instruct
📄 Paper · 🤗 Collection · 🌐 The Fin AI
Part of Open-FinLLMs — Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications (arXiv:2408.11878).
FinLLaMA-instruct (FinLLaMA-Instruct-8B in the paper) is FinLLaMA — LLaMA3-8B continually pre-trained on 52B financial tokens — instruction-tuned on 573K financial instructions to follow instructions and perform downstream financial tasks.
Model Details
| Base model | TheFinAI/FinLLaMA (LLaMA3-8B, continual pre-training) |
| Architecture | LlamaForCausalLM, 32 layers, hidden size 4096, ≈8.0B parameters, float16 |
| Instruction data | 573K samples after deduplication: FLUPE (123K), finred (32.67K), MathInstruct (262K), Sujet-Finance-Instruct-177k (177K) |
| Training | 8 × A100 80GB, ≈6 hours |
| Chat format | ChatML-style template (`< |
| License | Llama 3 Community License (inherited from the base model) |
Quick Start
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "TheFinAI/FinLLaMA-instruct"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Classify the sentiment of this headline as positive, negative or neutral: 'Company X beats quarterly earnings expectations.'"}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Evaluation
The paper reports that FinLLaMA-Instruct outperforms GPT-4 and other financial LLMs on 15 datasets. See the paper for results.
Intended Use & Limitations
- Research on financial instruction following and financial NLP tasks.
- Not investment advice; may hallucinate figures or facts.
Repository note
On 2026-10-07 this repository was restored to its LLaMA-8B state (commit 367dad6) after a Qwen2-0.5B checkpoint had been uploaded here by mistake in April 2025.
Citation
@misc{huang2025openfinllmsopenmultimodallarge,
title={Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications},
author={Jimin Huang and Mengxi Xiao and Dong Li and Zihao Jiang and Yuzhe Yang and Yifei Zhang and Lingfei Qian and Yan Wang and Xueqing Peng and Yang Ren and Ruoyu Xiang and Zhengyu Chen and Xiao Zhang and Yueru He and Weiguang Han and Shunian Chen and Lihang Shen and Daniel Kim and Yangyang Yu and Yupeng Cao and Zhiyang Deng and Haohang Li and Duanyu Feng and Yongfu Dai and VijayaSai Somasundaram and Peng Lu and Guojun Xiong and Zhiwei Liu and Zheheng Luo and Zhiyuan Yao and Ruey-Ling Weng and Meikang Qiu and Kaleb E Smith and Honghai Yu and Yanzhao Lai and Min Peng and Jian-Yun Nie and Jordan W. Suchow and Xiao-Yang Liu and Benyou Wang and Alejandro Lopez-Lira and Qianqian Xie and Sophia Ananiadou and Junichi Tsujii},
year={2025},
eprint={2408.11878},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2408.11878},
}
- Downloads last month
- 2
# Gated model: Login with a HF token with gated access permission hf auth login