Megadose AI progress, ranked and analyzed.

Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs

· HF Daily Papers ·
The paper proposes a black-box audit for hosted LLM APIs that measures fidelity loss without needing vendor logprobs.

Ventor-QTest sends repeated constrained prompts to reconstruct output distributions and report average fidelity loss. It also runs long-sequence probes to estimate extreme fidelity loss from upper-tail surprisal. In the reported tests, those fidelity measures varied by route and did not track GPQA-Diamond accuracy closely. The authors say extreme fidelity loss appeared alongside weaker Terminal-Bench pass rates as task exposure grew, suggesting long-horizon agent tasks may be more sensitive to it. HF Daily Papers' note

score 5

Categories: Research