Research and field notes on voice, data, and audio AI.
When is a Large Audio Language Model (LALM) judge a good enough proxy for human preference, and when do you still need a human in the loop?