Elektrine lite

← Feed

HalvarFlake

HalvarFlake@mastodon.social

<p>I do math. And was once asked by R. Morris Sr. : &quot;For whom?&quot;</p><p>Accidental two-time founder. Mathematician by education. Infosec luminary (has-been?).</p>

Posts

  • View post

    The EU gets a bad rap, but ... sometimes we have to appreciate that the 1949-2026 period is the longest period in *documented history* without interstate war among the bloodthirsty western European tribes.

  • View post

    I have a personal update: Next monday, I will be starting at @OpenAI working on better cyber (which will also entail some efficiency work). I&#39;m pretty excited about the things I will learn and the things we will do.

  • View post

    @patrick stimmt für closed source, bei FOSS projekten ist es subtilee

  • View post

    I feel very compute-constrained these days. In the past, I was &amp;quot;implementation time for experiment&amp;quot; constrained: I had an idea, but I needed to carefully weed out the ideas in order to focus on those I could implement for testing. Now that code agents allow faster experiment implementation, I am running out of compute all the time.

  • View post

    Ok, I have a 2048x3072 pixel video that shows the differences in training dynamics between Muon and AdamW, and they are fascinating to watch. Where can I upload/share such a video without it getting compressed to shreds?

  • View post

    Visualizing the decision boundaries and final result of the same problem optimized by Muon and by AdamW is fascinating. The solutions are qualitatively very different. AdamW has visibly redundant and useless neurons, Muon less so, but AdamW&amp;#39;s reconstruction looks more generic.

  • View post

    The weirdest observation: I generated movies visualizing the polytope boundaries for ReLU networks using Muon and AdamW. Same experiment, same data, same random seed. The difference is the &amp;quot;crease pattern&amp;quot; that the optimizers produce. The AdamW videos compress to 1/6th the size of the Muon videos. Something AdamW is doing allows the crease visualisation to be compressed well, but not Muon. This is the weirdest observation ever.

  • View post

    The weirdest observation: I generated movies visualizing the polytope boundaries for ReLU networks using Muon and AdamW. Same experiment, same data, same random seed. The difference is the &amp;quot;crease pattern&amp;quot; that the optimizers produce. The AdamW videos compress to 1/6th the size of the Muon videos. Something AdamW is doing allows the crease visualisation to be compressed well, but not Muon. This is the weirdest observation ever.

  • View post

    I don&amp;#39;t know what y&amp;#39;all are doing, but my experience with Fable and Sol is that they still need a lot of judgement, taste, guidance, and straight up correction to not mess up the codebases they operate in beyond repair.

  • View post

    In the last years, I wrote up some of the advice I often found myself giving to other founders, and a general list of lessons I learnt doing two companies, zynamics and optimyze. The full article - still work-in-progress - is here: https://thomasdullien.github.io/guides/entrepreneurship/

  • View post

    I made a small animation to illustrate how I think about simple ReLU-based feedforward networks.

  • View post

    https://www.faz.net/premium/digitalwirtschaft/thomas-dullien-zu-anthropics-mythos-software-war-nie-auf-perfekte-sicherheit-ausgelegt-das-raecht-sich-accg-200822228.html I wrote an FAZ guest article

  • View post

    ArXiv AI papers should list a &amp;quot;flops needed to reproduce experiments&amp;quot; in the abstract.

  • View post

    Slides for my QCon talk: https://docs.google.com/presentation/d/1wOT5kOWkQybVTHzB7uLXpU39ctYzXpOs2xVyD4zuYXY/edit?usp=drivesdk There&amp;#39;s a video on the QCon website, too, for virtual attendees.

  • View post

    I like going for a walk or a beer alone. Weirdly, feelings of boredom or loneliness are much more common for me when in the presence of others than when I am alone. Alone, I feel calm, and peace, and free.

  • View post

    This LLM security research from Anthropic also has important copyright implications: https://www.anthropic.com/research/small-samples-poison

  • View post

    I finally managed to write something about my recently deceased dear friend Felix &amp;#39;Fx&amp;#39; Lindner. https://phenoelit.de/fx.html#Halvar

  • View post

    I am digging through some PyTorch code, and ... I have to say: The Nvidia profiling infrastructure is very much no good. Terrible UX, terrible at answering the questions I need answered. Makes me pretty bullish about zymtrace.

  • View post

    LLMs are reshaping software dev. I don&amp;#39;t buy &amp;quot;the end of software dev&amp;quot;: Project ambition will grow dramatically. Ancient Egyptians could build the Pyramids but not the Empire State Building. Pre-LLM software will be viewed like we view the Pyramids.

  • View post

    @_dm killing bad ideas early.

  • View post

    Current events:

  • View post

    The zymtrace folks are killing it: https://zymtrace.com/article/anam-zymtrace/

  • View post

    I wrote some lines about mitigating vibe-coding risks by adopting a development model inspired by old-school computer breakin folks: https://addxorrol.blogspot.com/2026/03/slightly-safer-vibecoding-by-adopting.html

  • View post

    I am insanely proud about the following, even though I am not at all involved any more: https://opentelemetry.io/blog/2026/profiles-alpha/ Working with the optimyze team was so awesome.

  • View post

    I am seeing up to 30% wall-time fluctuation running the same code on the same data on the same VM type in GCE. That seems crazy to me. What&amp;#39;s the worst variation you&amp;#39;ve seen? How do you deal with these things? I mean, for any benchmarking you need to run the clickhouse tomato benchmark protocol? (E.g. run two instances of the same software on the same VM so they&amp;#39;re exposed to the same noisy neighbor effects).

  • View post

    Today&amp;#39;s insanity: