Reading the fine print, circling the suspicious bits...
Reading the fine print, circling the suspicious bits...
Shared read-only report
Underhyped: the paper claims “Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.” backed by BLEU score (28.4). The authors do contextualize this number themselves.
“Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train.”
“On the WMT 2014 English-to-German translation task, the big transformer model (Transformer (big) in Table 2) outperforms the best previously reported models (including ensembles) by more than 2.0 BLEU, establishing a new state-of-the-art BLEU score of 28.4.”