Safetensors
GGUF
English
gemma3_text
base
sml
void
pretrained-from-scratch
File size: 1,781 Bytes
c864e32
 
 
 
131fbc0
c864e32
 
 
 
 
 
ba77d40
b17fc0a
 
0c1aba0
 
 
 
6545f32
0c1aba0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b16fb7a
 
 
 
 
 
 
 
 
 
 
044dc84
8d17cba
7ed9b6d
8d17cba
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
datasets:
- CEAMFA/rewrite
- HuggingFaceFW/fineweb-edu
- appvoid/no-prompt-15k
language:
- en
tags:
- base
- sml
- void
- pretrained-from-scratch
---

<style>
  img {
  display: block;
  position: static;
  width: min(76%, 256px);
  height: auto;
  max-width: 100%;
  margin: 3rem auto 2.5rem;
  object-fit: contain;

  border: 2px solid rgba(255, 255, 255, 0.16);
  border-radius: 1rem;
  outline: none;

  user-select: none;
  -webkit-user-select: none;
  -moz-user-select: none;
  -webkit-user-drag: none;

  filter: none !important;
  transform: scale(1) !important;
  animation: none !important;
  box-shadow: none !important;

  position: relative;
  z-index: 1;
  transform-origin: center;
  }
</style>

<img src="https://huggingface.co/appvoid/void.0-preview/resolve/main/logo.png"/>

Introducing **void**: our first ever language model, trained from scratch with a novel hybrid tokenizer on 300M high-quality tokens (total of 2 epochs on a B300) using 4096 as context window. Total cost was $23 dollars. 132m parameters. Future releases are expected to be published in the following weeks/months.

| Benchmark     |   Accuracy | Normalized |
| ------------- | ---------: | ---------: |
| ARC Challenge |     25.17% |     27.22% |
| ARC Easy      |     45.03% |     43.01% |
| HellaSwag     |     31.57% |     35.77% |
| PIQA          |     61.15% |     60.12% |
| WinoGrande    |     53.12% |          — |
| ArithMark     |     35.20% |     35.20% |


If you want to sponsor future model releases, you can get information on how to make contributions here: [CEAMFA](https://huggingface.co/CEAMFA)

**Disclaimer:** Even though the model is based on gemma 3 architecture, the tokenizer is different so you might need to wait until this model can be added to llama.cpp