X / Twitter
i'm so excited, it looks like organizing a dataset by SAE activations can expose more structure than raw embeddings!
here are 2 umaps of 50k dad jokes embedded with nomic-text-v1.5
left is umap on the embeddings
right is umap of the top 64 features from SAE trained on FineWeb-edu-10BT sample
Gold. https://t.co/JGk2gPTT8P
Did I miss some breakthrough in vector quantization?
I need a gradient-descent friendly quantization. What is the best available solution?
you're definitely missing a lot if you haven't seen it... https://t.co/ctLGGulHcz
how have i never heard of @deepseek_ai?!
seems almost as good at coding as Claude 3.5 Sonnet, but 20x cheaper and open-source🤯