Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
rileyphone
on April 18, 2024
|
parent
|
context
|
favorite
| on:
Meta Llama 3
The bigger size is probably from the bigger vocabulary in the tokenizer. But most people are running this model quantized at least to 8 bits, and still reasonably down to 3-4 bpw.
kristianp
on April 19, 2024
[–]
> The bigger size is probably from the bigger vocabulary in the tokenizer.
How does that affect anything? It still uses 16 bit floats in the model doesn't it?
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: