
e=mc²9 min read
A Language Model as a Decision Model: Train a Tiny Transformer to Pick the Operation
Lesson 07: the Lesson 01 GPT, shrunk to 0.83M parameters, writes one token after a question: the operation that answers it. Beside it, pankhllm's hashed logistic regression. On phrasings neither saw, a threshold alone holds 98% precision for neither; with pankhllm's known-words guard, the transformer does, at 41.2% coverage.
2 views
Read