Reading the fine print, circling the suspicious bits...
Reading the fine print, circling the suspicious bits...
Shared read-only report
Underhyped: the paper claims “Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting.” backed by accuracy (86.4%). The authors do contextualize this number themselves.
“Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting.”
“When presented with examples formatted this way, GPT-3 achieves 86.4% accuracy in the few-shot setting, an increase of over 18% from the previous state-of-the-art.”