<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Pierre Peigné - Lefebvre]]></title><description><![CDATA[Pierre Peigné - Lefebvre]]></description><link>https://pierrepeignlefebvre.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!xbjB!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpierrepeignlefebvre.substack.com%2Fimg%2Fsubstack.png</url><title>Pierre Peigné - Lefebvre</title><link>https://pierrepeignlefebvre.substack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 30 Jul 2026 15:23:00 GMT</lastBuildDate><atom:link href="https://pierrepeignlefebvre.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Pierre Peigné - Lefebvre]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[pierrepeignlefebvre@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[pierrepeignlefebvre@substack.com]]></itunes:email><itunes:name><![CDATA[Pierre Peigné - Lefebvre]]></itunes:name></itunes:owner><itunes:author><![CDATA[Pierre Peigné - Lefebvre]]></itunes:author><googleplay:owner><![CDATA[pierrepeignlefebvre@substack.com]]></googleplay:owner><googleplay:email><![CDATA[pierrepeignlefebvre@substack.com]]></googleplay:email><googleplay:author><![CDATA[Pierre Peigné - Lefebvre]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Making benchmarks outputs directly useful for AI safety and security]]></title><description><![CDATA[Epistemic status: written in 30 min. This is not as polished as I&#8217;d like but I prefer share this as is than not to share it at all.]]></description><link>https://pierrepeignlefebvre.substack.com/p/making-benchmarks-outputs-directly</link><guid isPermaLink="false">https://pierrepeignlefebvre.substack.com/p/making-benchmarks-outputs-directly</guid><dc:creator><![CDATA[Pierre Peigné - Lefebvre]]></dc:creator><pubDate>Wed, 29 Jul 2026 12:19:16 GMT</pubDate><content:encoded><![CDATA[<p>AI are becoming increasingly good at solving problems. Benchmarks are saturating fast.</p><p>I think we should take advantage of this to make them solve useful problems while being evaluated. <a href="https://epoch.ai/frontiermath/open-problems">Epoch&#8217;s Open Problems</a> are already doing this, and this is great. It would be even better if the problems solved were directly relevant for AI safety or security. We can easily task AI to either optimize systems&#8217; performances (like the <a href="https://github.com/KellerJordan/modded-nanogpt">nanoGPT speedrun</a> but on code that is useful for safety this time) but also to find and fix vulnerabilities (in an adversarial red and blue team setup where you have both to break others&#8217; systems and to make sure yours is solid).</p><p>Examples of things that would be very useful to get</p><ul><li><p>Reducing the overhead of zero knowledge proof of training and inference (more on this very soon)</p></li><li><p><a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">Improving dramatically the robustness of sandboxes used to run and evaluate AI systems.</a> (Note: sandboxes will likely never be fully protected, it is just a matter of raising the bar as much as we can to save us time)</p></li><li><p><a href="https://www.theregister.com/security/2026/07/04/confidential-computings-core-trust-mechanism-is-broken-the-fix-may-not-exist/5266056">Improving dramatically the robustness of TEEs</a> (Note: same as above)</p></li><li><p>Optimizing code used to run safety experiments (i.e. getting the same results with less resources): for instance can you make this unlearning method as cheap as possible?</p></li><li><p>Improving code and methods of safety experiments (i.e. getting better results): can you make this unlearning method better/more robust?</p></li></ul><p>In all of such cases, leveraging a benchmark evaluation to actually get better tools could help significantly during the short window we still have.</p><p>I don&#8217;t have enough bandwidth to work on this myself but I&#8217;d love to see people taking care of this. Please reach out if you are interested.</p>]]></content:encoded></item></channel></rss>