You should watch the BBC video I’m linking below for a more in-depth explanation, but I’ll summarize here.
Your brain does not just take in information and relay it, it translates it into something that makes sense. And what you see trumps what you hear.
The perfect example of the McGurk effect in action is the scene in Greatest Showman where Jenny Lind sings “Never Enough”, which I am also linking below. Loren Allred sings the song, and Rebecca Ferguson acts it. In the song the word “never” repeats over and over again. During this section of the song, Allred is actually singing “neber”. Slightly changing consonants and vowel is a strategy singers use to more easily and consistently hit the notes. Rebecca Ferguson, however, always lip-syncs the word “never”. When the camera is away from her face you can hear the “b”, but when the camera is on her face, you can’t help but hear “v”.
https://www.youtube.com/watch?v=G-lN8vWm3m0
https://www.youtube.com/watch?v=kUkRoIMyqFo
https://www.youtube.com/watch?v=enrCBI7O_6I