MediaScout Weekly

  • Home
  • About
Sign in Subscribe

Latest

AgentLSD Rewrites How We Evaluate AI Security Agents

AgentLSD Rewrites How We Evaluate AI Security Agents

AgentLSD, listed at 2026-09-16 in the AgentSafety Papers tracker on GitHub, is a controlled framework for evaluating AI security agents against adversarial task contamination, a failure class most teams still have no name for. The framework runs those evaluations as CTF challenges. And the benchmark summaries report that it "
Bob M 17 Sep 2026
ComPO vs DPO: Tuning Llama-3-8B at 23GB Instead of 77GB

ComPO vs DPO: Tuning Llama-3-8B at 23GB Instead of 77GB

Bob M 17 Sep 2026
AI Coding Agents Didn't Get Adopted. They Got Defaulted.

AI Coding Agents Didn't Get Adopted. They Got Defaulted.

Bob M 16 Sep 2026
Open-Source AI Agent Frameworks Stopped Being Demos

Open-Source AI Agent Frameworks Stopped Being Demos

Bob M 16 Sep 2026
OpenAI Agents API Public Beta: Managed Runtime, Token-Only Billing

OpenAI Agents API Public Beta: Managed Runtime, Token-Only Billing

Bob M 15 Sep 2026
Open-Source LLM Evaluation Frameworks: What Actually Catches Agent Failures

Open-Source LLM Evaluation Frameworks: What Actually Catches Agent Failures

Bob M 15 Sep 2026
Open-Source Multi-Agent Orchestration Frameworks in 2026: The Lock-In Trap

Open-Source Multi-Agent Orchestration Frameworks in 2026: The Lock-In Trap

Bob M 14 Sep 2026
Show more

Subscribe to MediaScout Weekly

Don't miss out on the latest news. Sign up now to get access to the library of members-only articles.
  • Sign up
MediaScout Weekly © 2026. Powered by Ghost