2026/7/21 23:44:10 【Bug已解决】TPOTrainer.evaluate() returns NaN eval_loss while training loss is finite 解决方案
2026/7/21 23:44:10 【Bug已解决】GRPOTrainer environment_factory / tools is broken for VLMs whose tools return images 解决方案