<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[The Sim Gap]]></title><description><![CDATA[written by Fabian Friedland. Chief Strategy Officer Sim2Real Inc.]]></description><link>https://www.thesimgap.com</link><image><url>https://substackcdn.com/image/fetch/$s_!Vc9K!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b6e36c4-38a2-4053-9928-77e02f3f0d19_1200x1200.png</url><title>The Sim Gap</title><link>https://www.thesimgap.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 04 Aug 2026 17:42:11 GMT</lastBuildDate><atom:link href="https://www.thesimgap.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Fabian Friedland]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[fabianfriedland@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[fabianfriedland@substack.com]]></itunes:email><itunes:name><![CDATA[The Sim Gap]]></itunes:name></itunes:owner><itunes:author><![CDATA[The Sim Gap]]></itunes:author><googleplay:owner><![CDATA[fabianfriedland@substack.com]]></googleplay:owner><googleplay:email><![CDATA[fabianfriedland@substack.com]]></googleplay:email><googleplay:author><![CDATA[The Sim Gap]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[99.25% in Simulation Proves Nothing Yet]]></title><description><![CDATA[Four failures, and why every check we ran was blind to the same thing]]></description><link>https://www.thesimgap.com/p/9925-in-simulation-proves-nothing</link><guid isPermaLink="false">https://www.thesimgap.com/p/9925-in-simulation-proves-nothing</guid><dc:creator><![CDATA[The Sim Gap]]></dc:creator><pubDate>Mon, 03 Aug 2026 16:12:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4Hq_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>Zero Shot #2</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4Hq_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4Hq_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 424w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 848w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 1272w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4Hq_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png" width="1456" height="750" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:750,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59778,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://fabianfriedland.substack.com/i/209657472?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4Hq_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 424w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 848w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 1272w, https://substackcdn.com/image/fetch/$s_!4Hq_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc93387e9-7a88-42e3-b052-404d1a91d26b_1456x750.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesimgap.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>We train manipulation policies entirely in simulation. No human demonstrations, no real-robot training steps. A solver with full physics state generates the data; the policy only ever sees what a real robot&#8217;s cameras and encoders would see.</p><p>Current numbers. Single camera, behavioral cloning, 22k episodes: 94%. Two cameras, action chunking, 18k episodes: 99.25%.</p><p>Those go on the slide. The rest of this post is what it took, which was mostly a lot of runs that went nowhere.</p><div><hr></div><h2>Chunking on its own made things worse</h2><p>Action chunking is everywhere now. Predict a short sequence of actions instead of one, cuts down compounding error. Obvious win. We added it and success went to 28.5%.</p><p>Took me a day and several runs to accept the ablation was telling the truth. Chunking on its own is not a weaker version of the published result. It&#8217;s a different thing that performs badly, and the gap between the two is entirely in how you handle the predictions once you have them.</p><p>That part I&#8217;m not going to detail yet. What I&#8217;ll say is that the piece everyone treats as an implementation detail turned out to be the piece doing the work, and finding that out cost us about $150 of A10 time.</p><h2>Halving the resolution gave us exactly zero</h2><p>Not &#8220;degraded.&#8221; Zero. Every trial, failed.</p><p>256&#215;256 in works. 128&#215;128 in and the policy does nothing useful. There&#8217;s a floor where the thing you need to see stops subtending enough pixels, and below it there&#8217;s no signal to learn from. What it looked like in practice was a failure to identify the target at all, followed by directionless driving. Not a near miss, not a wrong choice between two objects. No evidence it had located anything.</p><p>We found this by ablation. We could have found it in twenty minutes with a ruler and a pixel count. I&#8217;ve since done the measurement properly, which is its own post.</p><h2>DAgger, twice</h2><p>DAgger is the textbook answer to distributional shift. Roll out the student, ask the expert what it should have done in the states the student actually reached, retrain, repeat.</p><p>It made us worse. Both times.</p><p>I ran it a second time assuming I&#8217;d broken something in the first attempt. That&#8217;s usually the right assumption. Same result, same direction.</p><p>I still don&#8217;t have a clean explanation. Working theory is that DAgger assumes an expert who can label any state the student reaches, and ours can&#8217;t. The privileged solver is very good along its own solution manifold and undefined on states it would never have entered itself. So when the student wanders somewhere bad and we query for a correction, we&#8217;re not asking an expert what to do. We&#8217;re asking an optimizer to operate outside its basin. It answers confidently. We train on the answer.</p><p>That&#8217;s testable and I haven&#8217;t tested it. If it&#8217;s right, the failure should get worse the more aggregation rounds you run, which matches what we saw, but so would half a dozen other explanations. What we did was stop running DAgger and put the compute somewhere else, which felt like giving up at the time.</p><h2>Validation loss lied to us for about a month</h2><p>This is the expensive one.</p><p>Lower val loss, worse driving. Repeatedly. More than once the checkpoint that won on loss lost badly in closed-loop eval, and not by a little.</p><p>Makes sense once you say it out loud. BC optimizes per-step action prediction. What you care about is whether a few hundred sequential decisions compose into a rollout that works. Those aren&#8217;t the same objective. But knowing that in the abstract and actually not reaching for the loss curve when you&#8217;re tired and picking a checkpoint turn out to be different skills.</p><p>Closed-loop eval is slower, noisier, annoying to automate. Use it anyway.</p><div><hr></div><h2>The thing they have in common</h2><p>I didn&#8217;t see it until I&#8217;d written them down next to each other.</p><p>In every case there was a check saying we were fine, and the check was built out of the same assumptions as the thing it was checking. Val loss is computed on data from the training distribution. Sim success is graded by the simulator that made the episodes. When the underlying assumption is off, the artifact and the check move together and the disagreement that would have told you never happens.</p><p>An oracle that shares an assumption with the thing it checks isn&#8217;t independent. It&#8217;s the same belief, wearing a different hat.</p><div><hr></div><h2>Not just a robotics thing</h2><p>I read a JuliaHub eval last month. Frontier LLMs deriving physics models from a spec, graded against sealed ground truth. Different field entirely, same failure, and honestly they described it better than I have.</p><p>One model ran twenty-two separate validation checks, the most in the study, and shipped an answer thirty percent wrong. All twenty-two verified its own equations against themselves.</p><p>Another built an independent reference implementation, which is the right instinct, then gave the reference the same truncated time horizon as the model it was checking. Blinded its own check.</p><p>The one that scored highest deliberately corrupted a field component to confirm the check could detect something known-bad. That&#8217;s the move. Costs nothing. Almost nobody does it, me included until recently.</p><div><hr></div><h2>Which brings me to our own problem</h2><p>Our teacher has privileged access: exact object poses, contact points, velocities, the lot. The student sees pixels and joint states. That asymmetry is the whole method.</p><p>But they&#8217;re both in MuJoCo.</p><p>Any place the simulator is wrong about the world, it&#8217;s wrong for both of them, identically. The student can match the teacher perfectly and score 99.25% on every sim metric we own and still be confidently wrong in exactly the shape MuJoCo is wrong. Our best result and our blind spot are the same number.</p><p>And by the argument above, no amount of additional sim-side evaluation catches it. They all inherit the assumption. I could run a thousand more eval episodes and learn nothing.</p><p>One unblinded oracle: the actual robot. It does the task or it doesn&#8217;t and nothing upstream gets a vote.</p><p>That&#8217;s what we&#8217;re running now, on a $50 RC car. Different platform to the arm the 99.25% came from, which is deliberate. If the method only transfers to the thing it was tuned on, it isn&#8217;t a method. I&#8217;ll post the number either way.</p><div><hr></div><p><em>The Sim Gap. Fabian Friedland, Sim2Real Inc.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesimgap.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The Deployment Problem Has a Simple Fix. Nobody Is Doing It. ]]></title><description><![CDATA[Introducing The Sim Gap &#8212; and the idea behind RealSIM]]></description><link>https://www.thesimgap.com/p/the-deployment-problem-has-a-simple</link><guid isPermaLink="false">https://www.thesimgap.com/p/the-deployment-problem-has-a-simple</guid><dc:creator><![CDATA[The Sim Gap]]></dc:creator><pubDate>Sun, 03 May 2026 19:36:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Vc9K!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3b6e36c4-38a2-4053-9928-77e02f3f0d19_1200x1200.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Zero Shot #1</p><p>Every robotics CTO I speak to has the same confession.</p><p>They&#8217;re spending 3 to 5 times more on deployment than they budgeted.</p><p>Not on hardware. Not on the model. On making the robot actually work for customers.</p><p>We call it the sim-to-real gap. It&#8217;s the industry&#8217;s open secret. We all know it exists. For manipulation tasks, we&#8217;ve largely given up on training in simulation, preferring to generate data on actual hardware using teleoperation.</p><p>---</p><p>What Everyone Gets Wrong</p><p>The standard response to the sim-to-real problem is: build a better simulator (until then, stick to humans and hardware).</p><p>Better physics. More accurate friction models. Higher-fidelity rendering. More domain randomization.</p><p>This is the wrong answer.</p><p>Not because better simulators aren&#8217;t useful &#8212; they are. NVIDIA&#8217;s Isaac Sim is excellent infrastructure and we build on top of it. But the problem isn&#8217;t simulator accuracy. The problem is that we&#8217;ve been training robots to be precise when we should be training them to be robust.</p><p>A policy that learned to exploit a specific friction coefficient? Dead on arrival in the real world.</p><p>A policy that learned smooth, margin-tolerant behavior that doesn&#8217;t depend on simulation quirks? Usually fine &#8212; even in a mediocre simulator.</p><p>The field has spent 20 years building better simulators. We think that&#8217;s the wrong hill.</p><p>---</p><p>The Training Wheels Insight</p><p>Here&#8217;s the question that led us to found Sim2Real:</p><p>What if the way we generate training data is the problem &#8212; not just the simulator itself?</p><p>Current best practice for training robot policies is expensive and unscalable. You need a robot that already knows how to do the task &#8212; via teleoperation, motion capture, or human demonstration &#8212; to generate the data you need to train a robot to do the task. It&#8217;s the classic catch-22.</p><p>Our approach is different. For tasks we want to teach a robot, we first solve the problem with full access to the internal state of the physics engine &#8212; exact object positions, ground truth everything. It&#8217;s allowed to cheat. We call this &#8220;GOD mode&#8221;. The solution may be a simple Python script for easy tasks, or a trained neural network for more difficult behaviors (this is not the final neural net).</p><p>Think of it as training wheels: the GOD mode solution can see things a real robot never could, and it uses that information to solve the task perfectly under myriad initial conditions.</p><p>Then we run this solution millions of times with randomized environments. Different object positions, different lighting, different friction, different starting states. Each run is logged &#8212; but critically, we only log what a real robot can actually sense. Camera feed. Joint angles. Proprioception. No physics engine access.</p><p>The dataset generated that way &#8212; robot-legal observations paired with correct actions &#8212; becomes the training set for a neural network that can solve the task with only data available to the real robot.</p><p>So now we remove the training wheels.</p><p>We train the neural net on this &#8220;clean&#8221; dataset, without the physics engine access we used to create the data. This neural net has to learn to perform the task using only what it can sense in real life. The gap between what the &#8220;cheating&#8221; solution knows and what the final neural net can see is precisely where robustness is learned.</p><div class="native-video-embed" data-component-name="VideoPlaceholder" data-attrs="{&quot;mediaUploadId&quot;:&quot;a0a86e3a-f86b-4257-a728-dd6cb0580565&quot;,&quot;duration&quot;:null}"></div><p>GOD mode solver: full physics access, perfect information. The deployed neural network never sees this data</p><p>Unlimited synthetic data. No human labeling. No teleoperation bottleneck.</p><p>---</p><p>Why This Works</p><p>The key insight is subtle but important.</p><p>When you give the robot privileged access to the physics engine, it finds the optimal solution for each randomized condition. When you strip that access from the neural net, it can&#8217;t rely on any specific physics assumptions &#8212; it has to generalize.</p><p>That forced generalization is deployment robustness.</p><p>The net doesn&#8217;t learn &#8220;pick up the cube when it&#8217;s at coordinates (x, y, z).&#8221; It learns &#8220;pick up the object that looks like this, from an approach direction like this, with a motion style that works across the range of conditions I was trained on.&#8221;</p><p>That&#8217;s a policy that survives contact with reality.</p><p>---</p><p>The Bigger Picture</p><p>The training wheels method is our near-term product. It solves the deployment problem today, using existing simulation infrastructure.</p><p>But we&#8217;re also working on something more ambitious.</p><p>Physics engines are, at best, an approximation. Even the best Newtonian engine we can build falls short of reality &#8212; because we are not God. The sim-to-real gap exists in part because no simulator perfectly captures the complexity of the real world.</p><p>What if the robot could define its own simulator?</p><p>Instead of building a better physics engine, let reality serve as the oracle. A correctly designed workflow could learn the dynamics that matter directly from real-world interaction, building a bespoke simulator tuned specifically to that robot in that environment. In the universe of all possible simulators, reality is the most accurate one.</p><p>That&#8217;s our moonshot. The training wheels method gets us to market. The learned simulator gets us to zero gap.</p><p>---</p><p>What We&#8217;re Building</p><p>This is the founding thesis of Sim2Real.</p><p>Three founders. Two of us &#8212; Dan Miller (CEO/CTO) and David Silver (COO) &#8212; previously co-founded On2 Technologies, which was acquired by Google. I&#8217;m Fabian Friedland, CSO, previously CEO of TychoBot where I spent two years deep in the sim-to-real problem before concluding the industry needed a fundamentally different approach.</p><p>We&#8217;re building the deployment layer for physical AI. Not a simulator &#8212; the methodology and tooling that makes simulation matter less.</p><p>---</p><p>The Sim Gap publishes weekly. The Zero Shot series goes deep on the technical foundations.</p><p>If you&#8217;re building robots, deploying models, or just tired of watching your policies fail on first contact with reality &#8212; subscribe.</p><p>---</p><p>Fabian Friedland is Chief Strategy Officer and co-founder of RealSIM (realsim.bot). He previously served as CEO of TychoBot and led Global Business Development at On2 Technologies (acquired by Google).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.thesimgap.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.thesimgap.com/subscribe?"><span>Subscribe now</span></a></p><h2></h2><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.thesimgap.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading The Sim Gap! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item></channel></rss>