πŸ”€ TinyLM-FC β€” Tiny Function-Calling Router

~0.47M Parameters | CPU < 4ms | MIT License

GitHub License: MIT Python 3.10+ Parameters CPU Latency GGUF

A sub-half-million-parameter decoder-only transformer trained completely from scratch to act as a lightweight, low-latency function-calling router. Given a user utterance, TinyLM decides whether to invoke an external tool (web_search) with structured arguments, or abstain (none) for chit-chat and greetings.

Built as an educational and empirical case study: how small can a specialized router be while still maintaining high precision?


πŸš€ Quick Links

Resource Link
πŸ“¦ Full Source Code, Training & Eval Pipeline github.com/Safeeq008/tiny-fc-lm
πŸ€— Model Weights (this page) huggingface.co/Safeeq/FCtiny
πŸ“„ Detailed README & Architecture Docs README.md on GitHub
🐳 Docker Deployment Dockerfile · DEPLOY.md
πŸ¦™ Ollama / llama.cpp Guide OLLAMA.md

πŸ“¦ Files in This Repository

File Size Description
model.safetensors ~1.8 MB PyTorch weights in safe zero-copy format
config.json ~1 KB Architecture hyperparameters
tokenizer.json ~121 KB Custom 2048-vocab ByteLevel BPE tokenizer
tinyfc.gguf ~1.0 MB Quantized model for Ollama & llama.cpp

⚑ Quickstart

Option A β€” Python (from this Hub)

git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm
pip install -r requirements.txt
python infer.py

Option B β€” GGUF via Ollama (zero-install inference)

# Download the GGUF and Modelfile from GitHub
curl -LO https://github.com/Safeeq008/tiny-fc-lm/raw/main/Modelfile

# Pull the GGUF from Hugging Face and register it
ollama pull hf.co/Safeeq/FCtiny:latest
ollama run hf.co/Safeeq/FCtiny

Or using the local GGUF:

# Download tinyfc.gguf from the Files tab above, then:
ollama create tinyfc -f Modelfile
ollama run tinyfc "weather in new york today"

Option C β€” FastAPI REST Server (from GitHub)

git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm && pip install -r requirements.txt
python serve.py 8000
  • Web UI: http://localhost:8000
  • Swagger docs: http://localhost:8000/docs

Option D β€” Docker

docker pull ghcr.io/safeeq008/tiny-fc-lm:latest  # or build locally:
git clone https://github.com/Safeeq008/tiny-fc-lm.git
docker build -t tiny-fc-lm . && docker run -p 8000:8000 tiny-fc-lm

πŸ” Python Inference (Manual)

import torch
from safetensors.torch import load_file
from tokenizers import Tokenizer

# ── Clone the repo for model.py ──
# git clone https://github.com/Safeeq008/tiny-fc-lm.git
import sys; sys.path.insert(0, "tiny-fc-lm")
from model import TinyLM

# 1. Load weights + tokenizer (download from Files tab above)
state_dict = load_file("model.safetensors")
tok = Tokenizer.from_file("tokenizer.json")

# 2. Instantiate model
model = TinyLM(vocab=2048, d=80, n_layers=4, n_heads=4, ffn_mult=4, max_len=80)
model.load_state_dict(state_dict)
model.eval()

# 3. Encode prompt with special tokens (<user>=1, <call>=2)
user_input = "weather in chennai today"
prompt_ids = [1] + tok.encode(user_input).ids + [2]

# 4. Greedily decode
with torch.no_grad():
    ids = torch.tensor([prompt_ids])
    for _ in range(30):
        logits = model(ids)[0]
        next_id = logits[0, -1].argmax().item()
        if next_id == 3:  # <end> token
            break
        ids = torch.cat([ids, torch.tensor([[next_id]])], dim=-1)
    output_ids = ids[0, len(prompt_ids):].tolist()
    print(tok.decode(output_ids))
# β†’ web_search|query=chennai weather today|recency=day

# For full constrained decoding + live DuckDuckGo dispatch, see:
# https://github.com/Safeeq008/tiny-fc-lm/blob/main/infer.py

πŸ“ Architecture

Hyperparameter Value Notes
Total Parameters 471,760 (~0.47M) Trainable weight count
Layers 4 Transformer decoder blocks
Hidden Dim 80 Embedding & layer dimensionality
Attention Heads 4 Head dim = 20 (even for RoPE)
Positional Encoding RoPE ΞΈ = 10000.0
Normalization RMSNorm Ξ΅ = 1e-6, pre-norm
FFN GELU (4Γ— width) Hidden dim = 320
Weight Tying βœ… Yes Input embeddings = output head
Vocabulary Size 2,048 Custom ByteLevel BPE
Context Length 80 tokens Max prompt + generation

πŸ“Š Benchmarks

Metric Val (In-Distribution) OOD (Held-Out Phrasings)
Exact Match Accuracy ~99.8% ~88.2%
Routing Decision Accuracy 99.9% 96.4%
Routing Precision (Tool) 99.9% 97.1%
Routing Recall (Tool) 99.9% 98.8%
Query Slot Exact Match 99.8% 89.5%
Recency Slot Accuracy 99.9% 97.2%
Syntactic Validity Rate 100.0% 99.8%
CPU Inference Latency ~3.2 ms ~3.4 ms

OOD = templates strictly held-out from training (unseen surface forms, same entity pools).


πŸ”€ Output Protocol

TinyLM outputs a strict, pipe-delimited schema:

web_search|query=<search query>|recency=<day|week|any>
none

Examples:

User Input Model Output
weather in tokyo today web_search|query=tokyo weather today|recency=day
latest AI news web_search|query=latest ai news|recency=day
iphone 16 price in india web_search|query=iphone 16 price india|recency=week
hi how are you none
tell me a joke none
what is 2 + 2 none

πŸ¦™ GGUF & Ollama

A tinyfc.gguf artifact (1.09 MB, fp16) is available in the Files tab above. This is a canonical Llama-architecture twin (512k params, SwiGLU MLP) trained on the same task and tokenizer, exported for llama.cpp-compatible runtimes.

Quick usage:

# Download tinyfc.gguf from the Files tab, then:
ollama create tinyfc -f Modelfile
ollama run tinyfc "latest ethereum price"
# β†’ web_search|query=ethereum price today|recency=day

See the full guide in OLLAMA.md.


πŸ—οΈ Reproduce from Scratch

Full pipeline available at github.com/Safeeq008/tiny-fc-lm:

git clone https://github.com/Safeeq008/tiny-fc-lm.git
cd tiny-fc-lm && pip install -r requirements.txt

python gen_data.py      # 1. Generate 80k synthetic training samples
python train_tok.py     # 2. Train custom 2048-vocab BPE tokenizer
python train.py         # 3. Train TinyLM-FC model (~5-10 min, CPU)
python eval.py          # 4. Full benchmark suite (val + OOD)
python infer.py         # 5. Interactive CLI with live DuckDuckGo search
python serve.py 8000    # 6. FastAPI server + web UI

# Optional: GGUF/Ollama path
python train_llama.py   # 7. Train Llama-arch twin
python export_gguf.py   # 8. Convert to GGUF
ollama create tinyfc -f Modelfile && ollama run tinyfc

⚠️ Known Limitations

  1. Single-tool: Strictly routes web_search only β€” no multi-tool schemas.
  2. English only: Custom BPE tokenizer trained on English synthetic corpus.
  3. 80-token context window: Inputs > ~60 user tokens are silently truncated.
  4. Synthetic training data: Real-world query distributions are far noisier; OOD generalization is limited to phrasing variation, not domain shift.
  5. Constrained decoding required: Raw greedy decoding may produce malformed outputs; use infer.py's constrained decoder for production.

πŸ“„ License

Released under the MIT License β€” free for research, benchmarking, commercial use, and edge deployment.

Repository: github.com/Safeeq008/tiny-fc-lm

Downloads last month
284
Safetensors
Model size
472k params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support