Section 1 : Hors podcast
- Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
- [2606.28186] Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
- Agentic Episodic Control
- [2510.09685] Deep Neural Networks Inspired by Differential Equations
- From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
- How Humans, Bots, and Agents Communicate About Vulnerabilities in Pull Requests
- When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
- [2606.28334] Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
- LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
- Generative AI Literacy Training Improves Intelligence Analysts’ Discrimination of Real and AI-Generated Images
- EMPATH: A Multilingual Auditor–Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
- Modelling Human Values for Value-Aware Multi-Agent Systems
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
- A Systems-Level Analysis of Sensitivity, Robustness, and Stability in Retrieval-Augmented Generation
- How Do LLMs Cite?
- Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
- Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare
- An AI agent for treatment reasoningover a biomedical tool universe
- LAMP: Lean-based Agentic framework with MCP and Proof Repair
- [2606.28856] Building AI-Ready Data Systems for Space Life Sciences, Aerospace Medicine, and Deep Space Exploration
- Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
- The strength of clinical evidence is recoverable from language model representations but not from their stated grades
- Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem
- Evidence-Informed LLM Beliefs for Continual Scientific Discovery
- Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
- MCP Server Architecture Patterns for LLM-Integrated Applications
- Attraction, Not Adaptation: How AI Agent Communities Develop Distinct Linguistic Identities
- Your Space is My Zone: Demystifying the Security Risks of AI-Powered Applications on Pre-Trained Model Hubs
- Proofs of Ownership for Machine Learning Models
- [2202.08832] Universality of empirical risk minimization
- Mutual Information Surprise: Rethinking Unexpectedness in Autonomous Systems
- Reinforcement Learning in Super Mario Bros: Curriculum, Pedagogy, and Optimal Level Design in World 1-1
- 1 Introduction
- Travel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge Graphs
- Shared Lexical Task Representations Explain Behavioral Variability In LLMs
- Emergent Culture in Minimal LLM Systems
- Investigating Multi-Agent Deliberation in Law
- Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?
- Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models
- Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits
- On the Convergence of Self-Improving Online LLM Alignment
- Automating Cause–Effect Specification with Knowledge Graphs and Large Language Models
- Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents
- Creating Intelligence
- A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control
- Unexplainability of Artificial Intelligence Judgments and Functional Implementation in Kant’s Perspective
- Learning AI Without a STEM Background: Mixed-Methods Evidence from a Diverse, Mixed-Cohort AIED Program
- How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis
- AI Assistance for Human Review of Default Judgments
- Taxing Artificial Intelligence
- AI usage patterns are shaped by perceived gains in human agency
- Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
- Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
- Black-Box Inference of LLM Architectural Properties with Restrictive API Access
- Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
- What Types of Human-AI Teams Exist?
- Controllable Sim Agents via Behavior Latents
- Artificial intelligence: Yann LeCun works on more flexible AI
- Agentics / Tech Things: Tokenmaxxing is dead, long live tokenmaxxing
- Companies Are Making Claude and Codex Talk Like Cavemen to Stop AI’s Soaring Costs
- Companies Are Throttling Employees’ AI Use Because It’s Too Expensive
- Software Engineering in the Age of AI
- Introducing Claude Sonnet 5 \ Anthropic
- Google’s new Nano Banana 2 Lite image model is its fastest and cheapest yet - Ars Technica
- Claude Sonnet 5 (max) - Intelligence, Performance & Price Analysis
- Issue 52 - Run Coding Models Locally, and Why Being Overqualified is a Risk • Buttondown
- Polaroid just dropped the most iconic anti-AI ad of the year
- Anthropic’s Claude Code Reportedly Uses Hidden Code to Detect Chinese Users
- Please stop the AI Confidence Theater - by Elena Verna
- Chain-of-Thought Spoofing Targets Reasoning AI Models
- MaralGPT “Mythos” 9B just released, and this is why I’m proud of my project. - A Beautiful Mind
- OpenClaw - L’assistant IA arrive sur iPhone et Android
- Claude Code planquait un mouchard dans la date du prompt
- DeepSeek V4 Launches in Mid-July with Peak-Valley Pricing
- Leanstral 1.5 - Mistral AI
- Why Won’t Europe Build AI Data Centers in Iceland?
- The real ROI of AI for knowledge work · Okane Land
- The Short Leash AI Coding Method For Beating Fable
- Lumo 2.0: The most powerful private AI
- Words Are a Byproduct of Consciousness. For LLMs, It’s Backwards.
- Alibaba to ban employees from using Anthropic’s coding tool, source says
- How working memory could give rise to consciousness
- AI has torched the market for junior programmers
- A quote from Jon Udell
- The AI Compass
- What’s new in Claude Sonnet 5
- Nano Banana 2 Lite
- Fable’s judgement
- A quote from Josh W. Comeau
- Open Source AI Gap Map
- Better Models: Worse Tools
- Cloudflare Pushes AI Companies To Pay For Publishers’ Content
- Ford rehires ‘gray beard’ engineers after AI falls short
- Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injection
- Microsoft builds a bouncer to keep bots out of Teams meetings
- Claude Sonnet 5.0 heads straight down the middle of the road to dodge controversy
- AI Can’t Be Listed as Inventor on Patent Applications, Japan’s Top Court Rules - The Japan News
- Agent Security Meets Regulatory Reality – A Practitioner Systematization of Autonomous-Agent Threats and Controls in Regulated Financial Systems
- Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations
- Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
- Robust Multi-Agent LLMs under Byzantine Faults
- Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
- On the Internet, Nobody Knows You’re an LLM Bot: Unmasking Web Agents with Multi-Layer Fingerprinting
- [2606.30360] On the Vulnerability of Parameter-Level Defenses to Model Merging
- Thinking Out Loud: Real-Time Deception Monitoring in Asymmetric LLM Negotiations
- AI Transparency: Governance Compliance or Stakeholder Requirements?
- https://emmettbuckthompson.com/blog/ai-is-making-us-lose-our-individuality
- Students Around the World are Using AI-Powered Smart Glasses to Cheat on Tests - Slashdot
- Claude Code Is Steganographically Marking Requests
- “AI Watermarking”: Bridging Policy Discourse and Technical Capabilities
- How Anthropomorphic Language Impacts Public Perceptions of AI
- Value–Action Alignment in Large Language Models under Privacy–Prosocial Conflict
- Decomposing Memorization Reduction in Privacy-Preserving Fine-Tuning of SLMs for CSIRTs
- Efficient Unlearning with Privacy Guarantees
Section 2 : Podcast 100% humain