Instructions to use Safeeq/FCtiny with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Safeeq/FCtiny with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Safeeq/FCtiny # Run inference directly in the terminal: llama cli -hf Safeeq/FCtiny
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Safeeq/FCtiny # Run inference directly in the terminal: llama cli -hf Safeeq/FCtiny
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Safeeq/FCtiny # Run inference directly in the terminal: ./llama-cli -hf Safeeq/FCtiny
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Safeeq/FCtiny # Run inference directly in the terminal: ./build/bin/llama-cli -hf Safeeq/FCtiny
Use Docker
docker model run hf.co/Safeeq/FCtiny
- LM Studio
- Jan
- vLLM
How to use Safeeq/FCtiny with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Safeeq/FCtiny" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Safeeq/FCtiny", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Safeeq/FCtiny
- Ollama
How to use Safeeq/FCtiny with Ollama:
ollama run hf.co/Safeeq/FCtiny
- Unsloth Desktop
- Docker Model Runner
How to use Safeeq/FCtiny with Docker Model Runner:
docker model run hf.co/Safeeq/FCtiny
- Lemonade
How to use Safeeq/FCtiny with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Safeeq/FCtiny
Run and chat with the model
lemonade run user.FCtiny-{{QUANT_TAG}}List all available models
lemonade list
- Atomic Chat
A sub-half-million-parameter decoder-only transformer trained completely from scratch to act as a lightweight, low-latency function-calling router. Given a user utterance, TinyLM decides whether to invoke an external tool (web_search) with structured arguments, or abstain (none) for chit-chat and greetings.
Built as an educational and empirical case study: how small can a specialized router be while still maintaining high precision?
π Quick Links
| Resource | Link |
|---|---|
| π¦ Full Source Code, Training & Eval Pipeline | github.com/Safeeq008/tiny-fc-lm |
| π€ Model Weights (this page) | huggingface.co/Safeeq/FCtiny |
| π Detailed README & Architecture Docs | README.md on GitHub |
| π³ Docker Deployment | Dockerfile Β· DEPLOY.md |
| π¦ Ollama / llama.cpp Guide | OLLAMA.md |
π¦ Files in This Repository
| File | Size | Description |
|---|---|---|
model.safetensors |
~1.8 MB | PyTorch weights in safe zero-copy format |
config.json |
~1 KB | Architecture hyperparameters |
tokenizer.json |
~121 KB | Custom 2048-vocab ByteLevel BPE tokenizer |
tinyfc.gguf |
~1.0 MB | Quantized model for Ollama & llama.cpp |
β‘ Quickstart
Option A β Python (from this Hub)
git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm
pip install -r requirements.txt
python infer.py
Option B β GGUF via Ollama (zero-install inference)
# Download the GGUF and Modelfile from GitHub
curl -LO https://github.com/Safeeq008/tiny-fc-lm/raw/main/Modelfile
# Pull the GGUF from Hugging Face and register it
ollama pull hf.co/Safeeq/FCtiny:latest
ollama run hf.co/Safeeq/FCtiny
Or using the local GGUF:
# Download tinyfc.gguf from the Files tab above, then:
ollama create tinyfc -f Modelfile
ollama run tinyfc "weather in new york today"
Option C β FastAPI REST Server (from GitHub)
git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm && pip install -r requirements.txt
python serve.py 8000
- Web UI:
http://localhost:8000 - Swagger docs:
http://localhost:8000/docs
Option D β Docker
docker pull ghcr.io/safeeq008/tiny-fc-lm:latest # or build locally:
git clone https://github.com/Safeeq008/tiny-fc-lm.git
docker build -t tiny-fc-lm . && docker run -p 8000:8000 tiny-fc-lm
π Python Inference (Manual)
import torch
from safetensors.torch import load_file
from tokenizers import Tokenizer
# ββ Clone the repo for model.py ββ
# git clone https://github.com/Safeeq008/tiny-fc-lm.git
import sys; sys.path.insert(0, "tiny-fc-lm")
from model import TinyLM
# 1. Load weights + tokenizer (download from Files tab above)
state_dict = load_file("model.safetensors")
tok = Tokenizer.from_file("tokenizer.json")
# 2. Instantiate model
model = TinyLM(vocab=2048, d=80, n_layers=4, n_heads=4, ffn_mult=4, max_len=80)
model.load_state_dict(state_dict)
model.eval()
# 3. Encode prompt with special tokens (<user>=1, <call>=2)
user_input = "weather in chennai today"
prompt_ids = [1] + tok.encode(user_input).ids + [2]
# 4. Greedily decode
with torch.no_grad():
ids = torch.tensor([prompt_ids])
for _ in range(30):
logits = model(ids)[0]
next_id = logits[0, -1].argmax().item()
if next_id == 3: # <end> token
break
ids = torch.cat([ids, torch.tensor([[next_id]])], dim=-1)
output_ids = ids[0, len(prompt_ids):].tolist()
print(tok.decode(output_ids))
# β web_search|query=chennai weather today|recency=day
# For full constrained decoding + live DuckDuckGo dispatch, see:
# https://github.com/Safeeq008/tiny-fc-lm/blob/main/infer.py
π Architecture
| Hyperparameter | Value | Notes |
|---|---|---|
| Total Parameters | 471,760 (~0.47M) | Trainable weight count |
| Layers | 4 | Transformer decoder blocks |
| Hidden Dim | 80 | Embedding & layer dimensionality |
| Attention Heads | 4 | Head dim = 20 (even for RoPE) |
| Positional Encoding | RoPE | ΞΈ = 10000.0 |
| Normalization | RMSNorm | Ξ΅ = 1e-6, pre-norm |
| FFN | GELU (4Γ width) | Hidden dim = 320 |
| Weight Tying | β Yes | Input embeddings = output head |
| Vocabulary Size | 2,048 | Custom ByteLevel BPE |
| Context Length | 80 tokens | Max prompt + generation |
π Benchmarks
| Metric | Val (In-Distribution) | OOD (Held-Out Phrasings) |
|---|---|---|
| Exact Match Accuracy | ~99.8% | ~88.2% |
| Routing Decision Accuracy | 99.9% | 96.4% |
| Routing Precision (Tool) | 99.9% | 97.1% |
| Routing Recall (Tool) | 99.9% | 98.8% |
| Query Slot Exact Match | 99.8% | 89.5% |
| Recency Slot Accuracy | 99.9% | 97.2% |
| Syntactic Validity Rate | 100.0% | 99.8% |
| CPU Inference Latency | ~3.2 ms | ~3.4 ms |
OOD = templates strictly held-out from training (unseen surface forms, same entity pools).
π Output Protocol
TinyLM outputs a strict, pipe-delimited schema:
web_search|query=<search query>|recency=<day|week|any>
none
Examples:
| User Input | Model Output |
|---|---|
weather in tokyo today |
web_search|query=tokyo weather today|recency=day |
latest AI news |
web_search|query=latest ai news|recency=day |
iphone 16 price in india |
web_search|query=iphone 16 price india|recency=week |
hi how are you |
none |
tell me a joke |
none |
what is 2 + 2 |
none |
π¦ GGUF & Ollama
A tinyfc.gguf artifact (1.09 MB, fp16) is available in the Files tab above. This is a canonical Llama-architecture twin (512k params, SwiGLU MLP) trained on the same task and tokenizer, exported for llama.cpp-compatible runtimes.
Quick usage:
# Download tinyfc.gguf from the Files tab, then:
ollama create tinyfc -f Modelfile
ollama run tinyfc "latest ethereum price"
# β web_search|query=ethereum price today|recency=day
See the full guide in OLLAMA.md.
ποΈ Reproduce from Scratch
Full pipeline available at github.com/Safeeq008/tiny-fc-lm:
git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm && pip install -r requirements.txt
python gen_data.py # 1. Generate 80k synthetic training samples
python train_tok.py # 2. Train custom 2048-vocab BPE tokenizer
python train.py # 3. Train TinyLM-FC model (~5-10 min, CPU)
python eval.py # 4. Full benchmark suite (val + OOD)
python infer.py # 5. Interactive CLI with live DuckDuckGo search
python serve.py 8000 # 6. FastAPI server + web UI
# Optional: GGUF/Ollama path
python train_llama.py # 7. Train Llama-arch twin
python export_gguf.py # 8. Convert to GGUF
ollama create tinyfc -f Modelfile && ollama run tinyfc
β οΈ Known Limitations
- Single-tool: Strictly routes
web_searchonly β no multi-tool schemas. - English only: Custom BPE tokenizer trained on English synthetic corpus.
- 80-token context window: Inputs > ~60 user tokens are silently truncated.
- Synthetic training data: Real-world query distributions are far noisier; OOD generalization is limited to phrasing variation, not domain shift.
- Constrained decoding required: Raw greedy decoding may produce malformed outputs; use
infer.py's constrained decoder for production.
π License
Released under the MIT License β free for research, benchmarking, commercial use, and edge deployment.
Repository: github.com/Safeeq008/tiny-fc-lm
- Downloads last month
- 284