I expected the larger Qwen3 model to show a clear advantage for writing correction. In this experiment, it didn't.
Across 20 paired writing cases, Qwen3 8B and 14B both achieved 19/20 complete-case corrections, while Qwen3 4B reached 18/20.
The dif...