Every scanned arXiv/bioRxiv paper is ranked here automatically, by hype gap.
Underhyped: the paper claims “Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting.” backed by accuracy (86.4%). The authors do contextualize this number themselves.
Underhyped: the paper claims “Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.” backed by BLEU score (28.4). The authors do contextualize this number themselves.