Reinforcement Learning · University of Zurich
RL Elicits Contextual Learning of Unseen Language Translation
RL with a chrF reward teaches an LLM to translate from in-context linguistic packets rather than memorize languages. On five unseen languages it averages 0.3335 chrF vs 0.2300 for SFT, which drops below the base model.