A model training method and apparatus is disclosed, where the model training method acquires a recognition result of a teacher model and a recognition result of a student model for an input sequence and trains the student model such that the recognition result of the teacher model and the recognition result of the student model are not distinguished from each other.