Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
๐ค
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
28
Follow
tyro12's profile picture
MariaIvanova's profile picture
KacLew's profile picture
103 followers
ยท
142 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
updated
a model
about 4 hours ago
AbstractPhil/mini-beatrix-2s
updated
a model
about 4 hours ago
AbstractPhil/mini-beatrix-1
replied
to
their
post
about 17 hours ago
12 day cook for mini beatrix v3 begins. This model's byte input is formatted using a method dubbed atlas input. ETA OCTOBER 2 2026 https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training/tree/main/mini-beatrix-3 https://github.com/AbstractEyes/geolip-bytelex https://github.com/AbstractEyes/alephllm Upgrades: * 46 billion byte training pipeline up from 16 billion * 32 block depth 376.0M in v3 up from 20 block 237.1M in 2s. * Active aleph head, repaired via the 2s faults and a large series of tests. * Byte atlas gateway router, explained below. * Guaranteed convergence follow-up AMOE arms on pretrain, fused into the final form, trained together over time to increase the collective capacity. * Multi-tokenizer oriented post-training arms distilled from multiple experts; E.G. Qwen 3.8 27b multi-layer teacher/student arms, CLIP big_g, Bert Code, and more. * Special token word implementation via AMOE arms is now tested up to 240 special tokens for routing. Theoretically each can implement it's own sub-arm aka nested commands. E.G; <think><think_symbolic> ... </think_symbolic></think> * Fused words post-training for faster inference. # The Atlas This atlas structure contains the conjoined shape of 12 tokenizers represented in the trigram format. This is used to predict difficulty in the overlaps, as per determined by the average byte overlap measured via corpus text and the compared overlap. This accuracy is only related to difficulty but it provides pre-training difficulty assessment that we will use to test post-training accuracy with it. This will determine if we can precalculate the likelihood of byte difficulty via tokenizer shape in byte form, for the multibyte fusion upcoming arm experiments for v3. The reason for this, is distillation. We need to train Beatrix to behave with multiple tokenizers, and this theory is showing accuracy with v1 and v2, but the 32 block depth of v3 will answer many questions alongside of the structure.
View all activity
Organizations
AbstractPhil
's buckets
1
Sort:ย Recently updated
AbstractPhil/alephllm-chat-storage
380 kB