Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🤝
Open to Collab
358.3
TFLOPS
AbstractPhila
PRO
AbstractPhil
19
6
27
Follow
3nesdeniz's profile picture
blkjack's profile picture
joonsoo-me's profile picture
94 followers
·
128 following
https://civitai.com/user/AbstractPhila
AbstractEyes
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
replied
to
their
post
about 7 hours ago
I believe I have a solution for cross-tokenizer chatter and noise, which I've built a prototype repo for this exact tooling dubbed bytelex. https://github.com/AbstractEyes/geolip-bytelex I had a bit of an inspiration recently and built a prototype for a token translation matrix that I called geolip-bytelex, which allows bytewise translation of many different tokenizers into byte format. The goal is to allow comparative distillation from multiple models to simultaneously represent expertise based on input tokens and differentiated teacher/student InfoNCE and MSE training paradigms, while cutting a huge cost of the distillation analysis comparative compute that cross-tokenizer noise will naturally cause when tokenizers are mismatched or incorrect, reducing a large portion of invalidity from the trained systems established by incorrect valuations from the distillations and lora trainings. Bytelex is essentially a byte-wise deconstruction of a tokenizer's state into a preliminary 255 byte language allowing for 10s of thousands of sequences per token to be represented rather than just a few. I'm not the first to try this, however I'm in a unique position due to my creation AlephLM being built entirely by learning it's own lexicon, thus allowing this to be more than experiment and instead a working prototype distillation potential. This can solve a longstanding multi-tokenizer problem that I and many other researchers have been facing, at the cost of setup overhead compute for the preliminary experiments, however the translation matrix I'm planning will potentially solve this problem allowing models to be directly bytewise captured in a more guaranteed methodology through cross-sampled analysis at distillation time in this optimizer state that I'm working out. I've dubbed this distillation loss ByteInfoNCE and the preliminary is showing humongous promise, with that the bytelex is the crux and prototype concept that I'll be expanding and researching further.
updated
a model
about 10 hours ago
AbstractPhil/geolip-bytelex
updated
a bucket
about 11 hours ago
AbstractPhil/alephllm-chat-storage
View all activity
Organizations
AbstractPhil
's datasets
83
Sort: Recently updated
AbstractPhil/alephllm-chat-history
Viewer
•
Updated
4 days ago
•
1
•
49
AbstractPhil/captionbert-8192-v2-consensus
Updated
17 days ago
•
242
AbstractPhil/conceptual-captions-12m-webdataset-berts
Viewer
•
Updated
18 days ago
•
32.3M
•
720
•
1
AbstractPhil/bulk-cc12m-features
Viewer
•
Updated
19 days ago
•
121M
•
3.88k
AbstractPhil/tower-probes-results
Viewer
•
Updated
28 days ago
•
17
•
174
AbstractPhil/qwen-deepfashion-fused
Viewer
•
Updated
Jul 12
•
122k
•
855
•
1
AbstractPhil/qwen-synth-characters-fused
Viewer
•
Updated
Jul 10
•
42.7k
•
413
AbstractPhil/qwen-synth-characters-100-json-test
Viewer
•
Updated
Jul 10
•
1k
•
54
AbstractPhil/anima-brent-90k-cache
Updated
Jul 5
•
77
AbstractPhil/qwen-synth-characters
Viewer
•
Updated
Jul 3
•
61k
•
127
AbstractPhil/qwen-deepfashion
Viewer
•
Updated
Jul 3
•
160k
•
217
AbstractPhil/diffusion-pipe-cache-test1
Viewer
•
Updated
Jun 27
•
8.92k
•
72
AbstractPhil/anima-90k-cache
Updated
Jun 26
•
40
AbstractPhil/diffusion-pretrain-set-ft1
Viewer
•
Updated
Jun 23
•
1.46M
•
1.98k
•
1
AbstractPhil/diffusion-pretrain-set-ft1-1024
Viewer
•
Updated
Jun 11
•
1.14M
•
642
AbstractPhil/sdxl-qwen-phase1-cache
Viewer
•
Updated
Jun 6
•
86k
•
191
AbstractPhil/geolip-sdxl-fid-scoring
Viewer
•
Updated
Jun 5
•
2.8k
•
55
AbstractPhil/sdxl-qwen-phase0
Viewer
•
Updated
Jun 4
•
86k
•
337
•
3
AbstractPhil/IMDB-PUBLIC-SCRAPED
Preview
•
Updated
May 19
•
135
•
1
AbstractPhil/ldhnam-deepfashion_controlnet
Viewer
•
Updated
May 19
•
26k
•
27
AbstractPhil/ffhq_flux_latents_repaired
Viewer
•
Updated
May 19
•
40.8k
•
199
AbstractPhil/synthetic-characters
Viewer
•
Updated
May 19
•
149k
•
340
AbstractPhil/CN_pose3D_V10_512
Viewer
•
Updated
May 19
•
66.5k
•
76
AbstractPhil/CN_pose3D_V7_512
Viewer
•
Updated
May 19
•
255k
•
80
AbstractPhil/synthetic-object-relations-json
Viewer
•
Updated
May 18
•
5k
•
17
AbstractPhil/cc-task1-json
Preview
•
Updated
May 18
•
61
AbstractPhil/cc-prompts-sharded
Viewer
•
Updated
May 15
•
3.32M
•
11
AbstractPhil/json-coco-format
Viewer
•
Updated
May 14
•
129k
•
153
AbstractPhil/svae-freckles-4096-cifar10
Viewer
•
Updated
Apr 10
•
60k
•
52
AbstractPhil/ryan-spearman-prepared-features
Viewer
•
Updated
Mar 27
•
1
•
84
Previous
1
2
3
Next