Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
lzhfdwu007's profile picture
muooon's profile picture
Felldude's profile picture
94 followers
·
127 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
updated
a model
about 3 hours ago
AbstractPhil/aleph-splat-0
updated
a model
about 9 hours ago
AbstractPhil/alephlm-0
replied
to
their
post
about 15 hours ago
The AlephLM results are rolling in and I'm very excited for the possibilities. I am very much looking forward to the coming weeks as I train the first AlephLM distillations from MANY teachers into AMOE arms. The AMOE arms hook cleanly to AlephLM structures and provide pos/neg learning elements. Hard positive and hard negatives coalesce to extend the capacity. https://huggingface.co/AbstractPhil/alephlm-0 https://huggingface.co/AbstractPhil/alephlm-adopt-0 As it stands they are structurally sound enough to fully pretrain. As or more stable than a standard Bert experimentally to distill using InfoNCE. AMOE legs improve these structures substantially. Structural behavior can be expanded in many ways on distilled and pretrained models alike. Attaching the AMOE to any model I've tried has created expanded or improved behavioral accumulations. They do have downsides but their upsides are very experimentally exciting. I've distilled multiple vits, multiple berts, and have begun distilling berts into AlephLM structures successfully. This is overall very exciting for me. I've begun formatting larger variants such as including GPT-2 and Qwen 3.5 4b as a paired combinator utilizing pathological T5 learned distilled encodings. It sounds odd, but the results show everything can be expanded and even be taught to cooperate. The CaptionBert-8192-v2 and v2-b are both structurally collapsing after token 480 or so, which is expected due to the small train. By distilling an AMOE arm to V2 by training with a longformer expert, the results are cutting through like butter. V2 has begun stabilizing rapidly for considerably longer token chains and sequences, the structure is repairing and building reusable capacity. I have discovered an improved methodology for sampling the AlephLM for text encoder benchmarks, which is predominantly L2 normalized outputs. Upcoming large paper for the distillation experiments and results within the next week or two. It's going to be a big one.
View all activity
Organizations
AbstractPhil
's datasets
82
Sort: Recently updated
AbstractPhil/captionbert-8192-v2-consensus
Updated
5 days ago
•
89
AbstractPhil/conceptual-captions-12m-webdataset-berts
Viewer
•
Updated
6 days ago
•
32.3M
•
570
•
1
AbstractPhil/bulk-cc12m-features
Viewer
•
Updated
7 days ago
•
121M
•
2.95k
AbstractPhil/tower-probes-results
Viewer
•
Updated
16 days ago
•
17
•
135
AbstractPhil/qwen-deepfashion-fused
Viewer
•
Updated
26 days ago
•
122k
•
3.22k
•
1
AbstractPhil/qwen-synth-characters-fused
Viewer
•
Updated
28 days ago
•
42.7k
•
2.43k
AbstractPhil/qwen-synth-characters-100-json-test
Viewer
•
Updated
28 days ago
•
1k
•
99
AbstractPhil/anima-brent-90k-cache
Updated
Jul 5
•
45
AbstractPhil/qwen-synth-characters
Viewer
•
Updated
Jul 3
•
61k
•
275
AbstractPhil/qwen-deepfashion
Viewer
•
Updated
Jul 3
•
160k
•
622
AbstractPhil/diffusion-pipe-cache-test1
Viewer
•
Updated
Jun 27
•
8.92k
•
82
AbstractPhil/anima-90k-cache
Updated
Jun 26
•
105
AbstractPhil/diffusion-pretrain-set-ft1
Viewer
•
Updated
Jun 23
•
1.46M
•
1.52k
•
1
AbstractPhil/diffusion-pretrain-set-ft1-1024
Viewer
•
Updated
Jun 11
•
1.14M
•
715
AbstractPhil/sdxl-qwen-phase1-cache
Viewer
•
Updated
Jun 6
•
86k
•
278
AbstractPhil/geolip-sdxl-fid-scoring
Viewer
•
Updated
Jun 5
•
2.8k
•
98
AbstractPhil/sdxl-qwen-phase0
Viewer
•
Updated
Jun 4
•
86k
•
228
•
3
AbstractPhil/IMDB-PUBLIC-SCRAPED
Preview
•
Updated
May 19
•
58
•
1
AbstractPhil/ldhnam-deepfashion_controlnet
Viewer
•
Updated
May 19
•
26k
•
24
AbstractPhil/ffhq_flux_latents_repaired
Viewer
•
Updated
May 19
•
40.8k
•
215
AbstractPhil/synthetic-characters
Viewer
•
Updated
May 19
•
149k
•
156
AbstractPhil/CN_pose3D_V10_512
Viewer
•
Updated
May 19
•
66.5k
•
66
AbstractPhil/CN_pose3D_V7_512
Viewer
•
Updated
May 19
•
255k
•
488
AbstractPhil/synthetic-object-relations-json
Viewer
•
Updated
May 18
•
5k
•
42
AbstractPhil/cc-task1-json
Preview
•
Updated
May 18
•
34
AbstractPhil/cc-prompts-sharded
Viewer
•
Updated
May 15
•
3.32M
•
9
AbstractPhil/json-coco-format
Viewer
•
Updated
May 14
•
129k
•
213
AbstractPhil/svae-freckles-4096-cifar10
Viewer
•
Updated
Apr 10
•
60k
•
54
AbstractPhil/ryan-spearman-prepared-features
Viewer
•
Updated
Mar 27
•
1
•
59
AbstractPhil/bertenstein-v1
Viewer
•
Updated
Mar 7
•
37.4k
•
1.21k
Previous
1
2
3
Next