I'm not "bitching" about it. Note my earlier reply to Robert:
"(Note that I'm not complaining here but, rather, truly wondering what the "science" behind any such "policy" must be)"
See above.
-----------------------------------------------------------^^^^^^^^^^^^^^^
You're making some (faulty) assumptions, here.
This isn't a "commercial" or a PSA. It's not something where you can sit down, listen to the N minutes of audio you laid down, decide *if* it is "good enough" and then depart. Then, come back next week to record something *different*, etc.
Note that I have explicitly mentioned the need to *analyze* the samples (see above quote), not "listen to" them.
E.g., one use is in building voices for speech synthesizers. The "recording session" is long and tedious. It is hard for a person to recite long lists of "nonsense words" while trying to maintain a "constant" speech pattern -- since the synthesizer may end up selecting a "unit" (phone, diphone, half-phone, etc.) from a "word" recorded at the start of the session and marrying it to a unit selected from a word recorded 15 minutes later in that same session. If the speaker's speech patterns have changed noticeably while reciting those words, you end up with crap.
But, you (me) don't even have access to those "units" in the recording studio! Instead, you have to crunch the data to identify and isolate them. Then, build the sample voice. Then, listen to *it* reciting real text to decide how good the set of units EXTRACTED FROM THE RECORDED AUDIO SAMPLE happen to be!
If the results aren't acceptable, you repeat the process.
If the "audio" results ARE acceptable but there is a lot of cuft in the signal (ambulance driving by), then you have to repeat the exercise. *Back* to the recording studio. Recite another/same set of nonsense words. Hopefully be able to recreate the exact physical and electrical environments. Then, *pray* that your speaker can mimic their earlier "performance" (lest you end up with the variation that I mentioned previously).
To put things in numerical perspective, English has ~50 phones. So, conceptually, a diphone-based synthesizer needs ~2500 diphones in its unit inventory. Granted, each "nonsense word" can contain more than one diphone (i.e., if you could get
5 *unique* diphones in each word, the speaker would only need to recite 500 words -- in the same manner, etc.). But, in practice, its not that easy.Other languages have more -- or less.
And, remember that you are selecting "speakers" based on the qualities of their natural voices. I.e., it may be an actor/actress hired just to provide speech samples. So, if they have to make repeated visits to this "studio" each time you uncover some problem with the previous sample set, it gets to be tedious -- on all involved.
[Imagine if *you* were that "voice model/actress" asked to return for the THIRD TIME to repeat more nonsense words... how much quality/effort are you likely to invest?]