Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
manh-linh's profile picture
Tico1982's profile picture
tyro12's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
published
an
article
about 1 hour ago
Raising Beatrix: A Byte-Level Model's Measured Childhood
updated
a model
about 2 hours ago
AbstractPhil/mini-beatrix-1
posted
an
update
about 3 hours ago
I believe I have a solution for cross-tokenizer chatter and noise, which I've built a prototype repo for this exact tooling dubbed bytelex. https://github.com/AbstractEyes/geolip-bytelex I had a bit of an inspiration recently and built a prototype for a token translation matrix that I called geolip-bytelex, which allows bytewise translation of many different tokenizers into byte format. The goal is to allow comparative distillation from multiple models to simultaneously represent expertise based on input tokens and differentiated teacher/student InfoNCE and MSE training paradigms, while cutting a huge cost of the distillation analysis comparative compute that cross-tokenizer noise will naturally cause when tokenizers are mismatched or incorrect, reducing a large portion of invalidity from the trained systems established by incorrect valuations from the distillations and lora trainings. Bytelex is essentially a byte-wise deconstruction of a tokenizer's state into a preliminary 255 byte language allowing for 10s of thousands of sequences per token to be represented rather than just a few. I'm not the first to try this, however I'm in a unique position due to my creation AlephLM being built entirely by learning it's own lexicon, thus allowing this to be more than experiment and instead a working prototype distillation potential. This can solve a longstanding multi-tokenizer problem that I and many other researchers have been facing, at the cost of setup overhead compute for the preliminary experiments, however the translation matrix I'm planning will potentially solve this problem allowing models to be directly bytewise captured in a more guaranteed methodology through cross-sampled analysis at distillation time in this optimizer state that I'm working out. I've dubbed this distillation loss ByteInfoNCE and the preliminary is showing humongous promise, with that the bytelex is the crux and prototype concept that I'll be expanding and researching further.
View all activity
Organizations
AbstractPhil
's datasets
83
Sort: Recently updated
AbstractPhil/alephllm-chat-history
Viewer
•
Updated
3 days ago
•
1
•
45
AbstractPhil/captionbert-8192-v2-consensus
Updated
16 days ago
•
240
AbstractPhil/conceptual-captions-12m-webdataset-berts
Viewer
•
Updated
16 days ago
•
32.3M
•
657
•
1
AbstractPhil/bulk-cc12m-features
Viewer
•
Updated
17 days ago
•
121M
•
3.88k
AbstractPhil/tower-probes-results
Viewer
•
Updated
27 days ago
•
17
•
145
AbstractPhil/qwen-deepfashion-fused
Viewer
•
Updated
Jul 12
•
122k
•
1.04k
•
1
AbstractPhil/qwen-synth-characters-fused
Viewer
•
Updated
Jul 10
•
42.7k
•
612
AbstractPhil/qwen-synth-characters-100-json-test
Viewer
•
Updated
Jul 10
•
1k
•
54
AbstractPhil/anima-brent-90k-cache
Updated
Jul 5
•
76
AbstractPhil/qwen-synth-characters
Viewer
•
Updated
Jul 3
•
61k
•
137
AbstractPhil/qwen-deepfashion
Viewer
•
Updated
Jul 3
•
160k
•
211
AbstractPhil/diffusion-pipe-cache-test1
Viewer
•
Updated
Jun 27
•
8.92k
•
67
AbstractPhil/anima-90k-cache
Updated
Jun 26
•
42
AbstractPhil/diffusion-pretrain-set-ft1
Viewer
•
Updated
Jun 23
•
1.46M
•
1.93k
•
1
AbstractPhil/diffusion-pretrain-set-ft1-1024
Viewer
•
Updated
Jun 11
•
1.14M
•
695
AbstractPhil/sdxl-qwen-phase1-cache
Viewer
•
Updated
Jun 6
•
86k
•
294
AbstractPhil/geolip-sdxl-fid-scoring
Viewer
•
Updated
Jun 5
•
2.8k
•
92
AbstractPhil/sdxl-qwen-phase0
Viewer
•
Updated
Jun 4
•
86k
•
322
•
3
AbstractPhil/IMDB-PUBLIC-SCRAPED
Preview
•
Updated
May 19
•
118
•
1
AbstractPhil/ldhnam-deepfashion_controlnet
Viewer
•
Updated
May 19
•
26k
•
26
AbstractPhil/ffhq_flux_latents_repaired
Viewer
•
Updated
May 19
•
40.8k
•
212
AbstractPhil/synthetic-characters
Viewer
•
Updated
May 19
•
149k
•
347
AbstractPhil/CN_pose3D_V10_512
Viewer
•
Updated
May 19
•
66.5k
•
69
AbstractPhil/CN_pose3D_V7_512
Viewer
•
Updated
May 19
•
255k
•
310
AbstractPhil/synthetic-object-relations-json
Viewer
•
Updated
May 18
•
5k
•
17
AbstractPhil/cc-task1-json
Preview
•
Updated
May 18
•
56
AbstractPhil/cc-prompts-sharded
Viewer
•
Updated
May 15
•
3.32M
•
11
AbstractPhil/json-coco-format
Viewer
•
Updated
May 14
•
129k
•
158
AbstractPhil/svae-freckles-4096-cifar10
Viewer
•
Updated
Apr 10
•
60k
•
56
AbstractPhil/ryan-spearman-prepared-features
Viewer
•
Updated
Mar 27
•
1
•
80
Previous
1
2
3
Next