Kim, M., Kim, C. W., & Ro, Y. M. (2023). Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video. Proceedings of the AAAI Conference on Artificial Intelligence, 37(7), 8273–8281. https://doi.org/10.1609/aaai.v37i7.25998