[1]
C. Zhang, “Aligning Language Models Using Follow-up Likelihood as Reward Signal”, AAAI, vol. 39, no. 24, pp. 25832–25841, Apr. 2025.