I think STE is to help humans examine model output especially its design, hence you can provide the right feedback
Yi Cui
AI & ML interests
Recent Activity
Organizations
I transformed a wordy repo:
Size
- Words: 115,559 โ 124,979 (+8.2%)
- Sentences: 4,795 โ 10,983 (+129%)
Sentence length
- Average words per sentence: 24.1 โ 11.4 (โ53%)
- 90th-percentile words per sentence: 46 โ 20 (โ57%)
- Longest sentence: 175 โ 62 words (โ65%)
- Sentences over 20 words: 53.1% โ 8.0% (โ85%)
- Sentences over 25 words: 41.6% โ 2.5% (โ94%
- Sentences over 40 words: 15.8% โ 0.2% (โ99%)
Paragraphs
- Paragraphs with more than 6 sentences: 1.2%
Readability
- Flesch reading ease: 54.1 โ 72.5 (+18 points)
- Flesch-Kincaid grade: 11.7 โ 6.0 (โ5.7 grad
Vocabulary and grammar
- Unique words: 7,233 โ 5,796 (โ20%)
- Passive voice per 1k words: 8.5 โ 5.7 (โ33%
- Complex tenses per 1k words: 1.5 โ 0.5 (โ66%)
- "-ing" words per 1k words: 26.9 โ 12.2 (โ55
Punctuation
- Semicolons per 1k words: 9.8 โ 0.9 (โ91%)
- Em dashes per 1k words: 8.9 โ 0.1 (โ99%)
Can't fine-tune or prompt engineer this model so you have to accept its accuracy as is. Add other harnesses if the accuracy is not as good
Really no idea. Just wait for the open source ones to come out
1. Very likely a small model. You can certainly pretrain, but I would grab an existing base model, say Qwen 3 class
2. The new RL method is a breakthrough, classification doesn't need to align with human preferences
3. The new output is an overstatement. It's just a new LM head. Of course autoregressive decoding can be used for classification: it takes just a few tokens to express the output. Think twice: are you sure classification doesn't need few-shot, CoT, or reasoning? All of these depend on auto-regressiveness
4. It carves out a market already existing, which is now served by oversized LLMs (hence overpaid), e.g. LLM as judge, labeling
5. Jevons effect will kick in, promoting more modeling efforts for small budget teams. It might even accelerate RSI
๐คฃ
Astra: ๐ง ๐ง ๐ง ๐ฌ
Really appreciate everyone's knowledge sharing and critique here.
Isn't this a gloomy picture that no frontier model can be run on the edge? Smaller models are for workflows and automations.
We are all used to conversing with a frontier model and let it run things on your device. A weaker model will degrade your current experience hence no adoption.
Wow! thx for all the works here. I wish more people can see the methodology.
My Mac is only 16GB ๐
We need to find one of those people above who have big budget
https://www.ebay.com/itm/257712960417
MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K.
On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX.
For high-end inferencing, I place my bet on workstations over Macs.
For railroad, it was the track, and dot com, fiber. Compared to these two, GPUs depreciate too fast.
The answer is power plants.
Method 1: Assuming Ox Alpha is the GLM 5.3 class, it has roughly the same active param count as DeepSeek V4 Pro (~40B vs 49B), derive the cost by the floor price DS has ever published.
Method 2: Assuming the users are concentrated within an 8-hour working window each day, derive number of H800 nodes needed (~700 nodes at $2/GPU-hour).
In both methods, I assume 90/10 IO split and 60% cache hit. Both methods come to $2M.
I think the ROI is awesome: (1) publicity and (2) data harvesting.
Whoever this is, they give away so much compute for a model launch
Most likely
Lots of SLMs, like the 27b one from Qwen