LLM-as-a-Judge: The Importance of Harshness
his piece tackles a problem that’s becoming more relevant as AI tools spread into everyday work: how do you check the quality of something an AI produced, when there’s no answer key to grade against?
his piece tackles a problem that’s becoming more relevant as AI tools spread into everyday work: how do you check the quality of something an AI produced, when there’s no answer key to grade against?
Vandalizer 4.8 has just released: Here are the major changes Your Knowledge Bases Are More Reliable PDFs now upload and process correctly every time — no more silent failures. Answers also come with clickable citations so you can see exactly where they came from. And if a question isn’t covered by your documents, the assistant…
Vandalizer v4.5.0 Released
July 9, 2026
11am PDT
Technical questions for Vandalizer
Research administration is risk-averse on purpose. The work navigates federal regulations, sponsor-specific terms, institutional policies, and audit trails that survive personnel changes by years. RAs have been trained well, and the training has stuck: when in doubt, slow down; when the rule is unclear, ask; when the answer is unverifiable, do not submit. That posture is not a deficit to be retrained. It is an asset the institution already has, and it should be respected as such.
Vandalizer v4.5.0 Released
In this interview we discuss how Research Administration offices vary widely in scale and mission. Building Vandalizer—an AI tool suite flexible enough to serve them all—has required diverse institutional perspectives from the start.
Have questions about Vandalizer, research data, or AI? Come join us for monthly office hours!
When: Fourth Tuesday of every month at 11:00am PDT
Where: Zoom
The blog post explains the shift from prompt engineering, carefully wording individual requests to get better AI responses, to context engineering, which focuses on designing the entire information environment an AI uses to produce results.
by Michael Overton, PhD Original post