I agree, but I also really think it’s worth sticking to zarf’s point here. The difference between big studio and indie dev, or paid and volunteer voice work, is basically irrelevant for many of our games – because the combinatorial complexity of those games’ text makes it realistically impossible for anyone (rich/poor, contractor/volunteer) to do a human-read audio version of them in the first place.
I write in ChoiceScript, not parser games, but that doesn’t make VO as simple as reading out a bunch of text blocks. My first game was hundreds of thousands of words long, and had a lot of in-line variation throughout, based on things like the character’s social class, gender, relationship with the other characters, and past actions.
My sequel-in-progress doubles down (God help me) on both the length and the complexity, with lots of variation based on who’s accompanying you in a given scene. (Including plenty of cases where whoever happens to be there delivers the same line – which is a little more efficient in terms of text, but might easily quintuple the voice acting demand for the scene.)
Those games can be screen-read by machine, without (at present) much in the way of feeling or nuance or vocal variation. To get them narrated by humans (without dumping the variety) would take way more hours than anyone could ever justify spending on voice acting. Even if you just focused on the dialogue, and cut features like the reader’s ability to input a name of their choice.
Obviously not all text-based IF has that kind of variability…but a whole lot of it crosses the complexity threshold where human voice acting would be prohibitively time-expensive (especially if you had to justify the cost against the expected income from your audience…but honestly, I don’t think we’re in remotely reasonable labor-of-love volunteer territory for much of it either).
And that’s relevant to @jkj_yuio’s original question. If we take AI voice generation off the table, we’re not (in these cases) providing work for human voice actors; we’re just enforcing the practical impossibility of a voiced version.