SPEECH PERCEPTION
Fromkin and Rodman (1988) provide the simplest explanation for speech perception. They argue that when you hear a car backfire, you may wonder whether the sound represented a gunshot or a backfiring. Your perception of the particular acoustic signal and your knowledge of what creates different sounds result in your assigning some "meaning" to the sounds you heard. Similarly, when you hear the sounds represented by the phonetic transcription [ʤæk ɪz ən ˈi:dɪət], you assign the meaning "Jack is an idiot" to the sound signal. The acoustic signal, however, does not reach our ears in phonemic segment-sized chunks; it is a semicontinuous signal. In order for us to process it as speech, it must be segmented into phonemes, words, phrases, and sentences.
Speech perception is a process by which we segment the continuous signal and, in so doing, may "mischunk" or misperceive the speaker's intended utterance. The difficulties inherent in speech perception are compounded by the fact that the speech signal for the "same" utterance varies greatly from speaker to speaker and from one time to the next by the same speaker. Nevertheless the brain is able to analyze these different signals, conclude that they are the same linguistically, segment the utterance into a phonetic/phonological string of words, "look up" the meaning of these words in the mental dictionary, technically called the lexicon, analyze the linear string of words into a hierarchical syntactic structure, and, most of the time, end up with the intended meaning.
All this work is done so quickly we are unaware that it is going on at all. Despite the variation between speakers and occurrences, there must be certain invariant features of speech sounds that permit us to perceive a /d/ or an /ə/ produced by one speaker as identical phonologically with a /d/ and /ə/ produced by another. The relations between the formants (i.e., a frequency range where vowel sounds are at their most distinctive and characteristic pitch) of the vowels of one speaker are similar to those of another speaker of the same language, even though the absolute frequencies may differ.
When a stop consonant is produced, the signal is interrupted slightly, and the frequency of the "explosion" that occurs at the release of the articulators in producing stop consonants differs from one consonant to another. The transitions between consonants and vowels provide important information as to the identity of the consonants. After voiceless consonants, the onset of vowel formants starts at higher frequencies than after voiced consonants. Different places of articulation influence the starting frequencies of formant onsets.
There are many such acoustic cues that, together with our knowledge of the language we are listening to, permit us to perform a "phonetic analysis" on the incoming acoustic signal. Confusions may also be disambiguated by visual, lexical, syntactic, and semantic cues. Speech communication often occurs in a noisy environment, but we can still pick out of the sound signal those aspects that pertain to speech. We are thus able to ignore large parts of the acoustic signal in the process of speech perception, which has led to the view that the human auditory system—perhaps in the course of evolution—has developed a special ability to detect and process speech cues.
To understand an utterance we must, in some fashion, retrieve information about the words in that utterance, discover the structural relationship and semantic properties of those words, and interpret these in the light of the various pragmatic and discourse constraints operating at the time. Further, all of this takes place at a remarkably rapid pace. Analyzing the speech signal in speech perception is a necessary but not sufficient step in understanding a sentence or utterance. Suppose you heard someone say:

You would still be unable to assign a meaning to the sounds, because the meaning of a sentence depends on the meanings of its words, and the only English lexical items in this string are the morphemes a, is, and -ing. The sentence lacks any English content words. You can only know that the sentence has no meaning by attempting a lexical lookup of the phonological strings you construct; finding no entries for sniggle, blick, prock, or slar in your mental dictionary tells you that the sentence is composed of nonsense strings. If instead you heard someone say
The cat chased the rat
through a lexical lookup process you would conclude that an event concerning a cat, a rat, and the activity of chasing had occurred. Who chased whom is determined by syntactic processing. That is, processing speech to get at the meaning of what is said requires syntactic analysis as well as knowledge of lexical semantics. Stress and intonation provide some cues to syntactic structure. We know, for example, that the different meanings of the sentences

can be signaled by differences in their stress patterns. Relative loudness, pitch, and duration of syllables thus provide important information in speech perception.