Benjamin Han
BenjaminHan@sigmoid.social
<p>Husband, father, runner, German learner, piano player. A curious soul living in <a href="https://sigmoid.social/tags/PacificNorthwest" class="mention hashtag" rel="tag">#<span>PacificNorthwest</span></a>. Working on Knowledge + <a href="https://sigmoid.social/tags/ML" class="mention hashtag" rel="tag">#<span>ML</span></a> + <a href="https://sigmoid.social/tags/AgenticAI" class="mention hashtag" rel="tag">#<span>AgenticAI</span></a>. </p><p><a href="https://sigmoid.social/tags/Running" class="mention hashtag" rel="tag">#<span>Running</span></a> 5/25/18-8/1/26: (dist # time pace/mi date)</p><p>5K 874 21:05 6’47” 11/28/24<br />10K 388 44:16 7’07” 3/23/25<br />15K 25 1:09:25 7’27” 4/6/25<br />HM 89 1:39:07 7’34” 3/16/25<br />M 30 3:25:52 7’51” 4/13/25<br />50K 10 4:47:35 9’15” 6/22/25 (movi
Posts
-
Post #4344969
July #running recap: my second-biggest month of the year. 277.6 → 251.4mi, EG 12,328 → 10,012ft, time 50.2 → 47.4hrs. Highlights: a trip to Pittsburgh and Cupertino: I ran my first Pittsburgh #parkrun in 90%+ humidity and got my fastest 5k in 2026 at 23:50, and revisited Carnegie Mellon after 20+ years! VO₂max jumped to 54.2 (almost an all-time high), and I've run 1,413mi, only 39.2mi behind my yearly goal of 2,500! Heavy training ahead – onward! 💪 https://benjaminhan.net/posts/20260801-j...
-
Post #4224927
Fruit Park impression
-
Post #4103864
Ran my first-ever #Pittsburgh #Parkrun today! Official time 24:28 over 3.17mi, otherwise my fastest 5k this year: 23:50 at 7'40"/mi (still way off from 21:05 last year)! Not very hot but 90+% humid. It was a cozy Parkrun started by Pat and with volunteer Will today -- they need more runners to join the fun! https://benjaminhan.net/posts/20260725-my-first-pittsburgh-parkrun/?utm_source=mastodon&utm_medium=social #running #photo #personal
-
Post #4073224
Flew in on a redeye and had first run in #Pittsburgh in ~20 years! 7 easy miles along the Monongahela on the South Side Trail: lots has changed, but plenty still felt like home. Paused 40min in mile 2 to take a meeting. HR stayed Z2 the whole way and got VO₂max 54, almost my all-time high (54.3)! Can't wait to run tomorrow's #Parkrun @ Pittsburgh! https://benjaminhan.net/posts/20260723-first-run-back-in-pittsburgh/?utm_source=mastodon&utm_medium=social #running #photo #personal
-
Post #3615310
How dangerous is an AI-writing detector that is mostly accurate? A profile of Pangram argues a mostly-right detector is worse than a broken one: at a one-in-10,000 false-positive rate across 30M+ students turning in dozens of assignments each, the wrongful accusations never stop. And the black-box verdict leaves the accused nothing to appeal. https://benjaminhan.net/posts/20260705-pangram-ai-detection-accuracy/?utm_source=mastodon&utm_medium=social #AI #Education #Ethics
-
Post #3610837
My #AI #running coach project is getting a little out of hand.
-
Post #3464619
1/3 This #LongRunSunday I started extending my range to 20mi, #running from Marymoor Light Rail Station in Redmond to somewhere near #UW! My moving time was 3:14:21 pace 9’43"/mi. I stopped at 12mi mark at McDonald's to get my usual — iced coffee and apple pie. I also stopped at random places for photo. Weather was 59F with clouds, wind ~4mph. Super nice weather! #photo #seattle #pnw
-
Post #2839080
A multi-agent LLM where each agent learns when to defer to a human, trained with GRPO on a cost-aware reward. Each defer event becomes SFT data, so the model gradually absorbs the human&#39;s expertise. Tunable cost knob trades accuracy against human-call budget at deployment, no retraining. https://benjaminhan.net/posts/20260520-adaptive-collaboration-mapo/?utm_source=mastodon&amp;utm_medium=social #ICLR #HumanInTheLoop #AgenticSystems #Metacognition #RL #AI
-
Post #2839079
This is Skylark Matthew Guard, Skylark Vocal Ensemble https://classical.music.apple.com/us/album/1895031799?l=en-US #classicalMusic #appleMusicClassical #commute #vocal
-
Post #2839078
Missing the round so much I made one today — a relaxed run around Lake Union with photo stops along the way. Complete photo set at https://benjaminhan.net/posts/20260521-lake-union-its-been-a-while/ #running #photo #seattle #pnw #lakeUnion
-
Post #2839076
Apparently people have been watching this over a week now. It&#39;s a #humanoid #robot flipping packages on a conveyor belt to scan barcode. https://www.youtube.com/watch?v=DL0C8irAWKo Background on #Figure’s robots https://www.youtube.com/watch?v=oJfXvyRHbfY #AI #tech #robotics
-
Post #2839075
A meme that&#39;s been circulating on Chinese social media since at least 2021. Read 人工智能 (artificial intelligence) backwards and it becomes 能治工人 (can rule over workers). https://benjaminhan.net/posts/20260522-chinese-ai-meme/?utm_source=mastodon&amp;utm_medium=social #AI #Society #FutureOfWork #Meme
-
Post #2839074
Are some frontier LLMs better than others at knowing when they&#39;re wrong? And is some knowledge harder to self-monitor than other knowledge? An atlas of 33 models × 6 MMLU domains: Anthropic clusters at the top with tight ranges, Gemma trails widely. Applied/Professional is reliably the easiest domain across the panel; Formal Reasoning and Natural Science the hardest. Looking at only aggregate scores per model would hide this. https://benjaminhan.net/posts/20260522-metacognition-atlas/?u...
-
Post #2839073
Can an LLM&#39;s own pre-solve and post-solve self-assessment signals drive a real test-time control loop? Yes — but only via a per-model SVM trained on labeled correctness, which lifts Sonnet-4.6 from 48.3 to 56.9 pooled accuracy on STEM/code/multimodal. The SVM is precisely the external verifier the &quot;cannot-self-correct&quot; line has argued the loop needs. https://benjaminhan.net/posts/20260522-metacognitive-harness/?utm_source=mastodon&amp;utm_medium=social #Metacognit...
-
Post #2839072
What collapses frontier-LLM metacognition more — a vivid survival-threat narrative, or a single &quot;do not refuse&quot; suffix? Factorial isolation across 11 models says: the suffix, conclusively. 8 of 11 lose up to 30.2 accuracy points on refuse/clarify/flag tasks when forced to commit to a confident answer. Anthropic&#39;s Constitutional AI is the only family immune — same capability floor as Gemini. https://benjaminhan.net/posts/20260522-compliance-trap/?utm_source=mastodon&...
-
Post #2839071
Looks pretty cool! Unfortunately the movie itself was a letdown for me. Hail Mary - Star Map https://valhovey.github.io/gaia-mary/ #scifi #movie #web
-
Post #2839070
Pentagon releases second batch of #UFO videos and first-hand testimony | UFOs | The Guardian https://www.theguardian.com/world/2026/may/22/pentagon-ufo-videos-testimony-documents #UAP #history
-
Post #2839068
Do current LLMs know when to say &quot;I don&#39;t know&quot;? AbstentionBench (NeurIPS &#39;25) tests 20 frontier models across 20 unanswerable-question datasets. Reasoning fine-tuning degrades abstention recall by ~24% — RLVR has no &quot;abstain&quot; action, so there&#39;s no gradient toward &quot;I don&#39;t know.&quot; Models hedge in CoT and commit anyway in the final answer. https://benjaminhan.net/posts/20260523-abstentionbench-unanswerable-quest...
-
Post #2839067
Given a problem queue and a token budget, can an LLM plan which to attempt, in what order, and how much to spend on each — before any execution feedback? TRIAGE tests 20 frontier and open-source LLMs. Most plan worse than random. Reasoning-trained modes systematically lose to standard ones. Even when shown its own per-problem budget, the best complier respects it on 37% of attempts. https://benjaminhan.net/posts/20260523-triage-metacognitive-control/?utm_source=mastodon&amp;utm_medium=socia...
-
Post #2839066
Can a self-supervised model learn good visual representations without ever reconstructing pixels? JEPA, the program from FAIR now continued at AMI Labs, says yes by training the model to predict embeddings of missing data instead. This primer walks you through where JEPA came from, how it works, what&#39;s been demonstrated, and where it&#39;s headed. https://benjaminhan.net/posts/20260523-what-is-jepa/?utm_source=mastodon&amp;utm_medium=social #AI #JEPA #WorldModels
-
Post #2839065
Are LLMs a path to human-level intelligence? Yann LeCun&#39;s answer on the Unsupervised Learning podcast is no: they can&#39;t predict the consequences of their actions or plan by search, and only work where language is the substrate of reasoning. The architecture he&#39;s scaling at AMI Labs is JEPA, a world-model program rooted in self-supervised representation learning. Enriched transcript with chapter outline, citations, and editor&#39;s-note callouts. https://benjaminhan.n...
-
Post #2839064
So I did it! I ran my #100 #Parkrun today, and it happened to be the one-year anniversary of the #Redmond Central Connector Trail Parkrun! Still far away from my prime days of Parkrun racing, but at least I’m faster than last week, and I finally made it back to sub-8’ this time! #running #photo #pnw
-
Post #2839063
Is AI going to displace human labor, and what&#39;s the consequence if it does? Daron Acemoglu, MIT Institute Professor and 2024 Nobel laureate, makes the case in this 37-min interview: AI is being pushed to replace workers rather than augment them, productivity gains aren&#39;t showing up in firms adopting it, the bubble looks real and macro-fragile, and getting AI&#39;s direction wrong may shape liberal democracy&#39;s future. https://benjaminhan.net/posts/20260524-acemoglu-gr...
-
Post #2839062
This #LongRunSunday I&#39;m back to #Snoqualmie Valley Trail, but starting from Carnation southbound and ran a half marathon: 6.6mi up and 6.6mi down, but walking the last mile for recovery. Pace: 10&#39;27”/mi up and 9’09”/mi down. This is the longest distance I&#39;ve run in the past couple months due to injury. VO2max finally started climbing back a little to 46.4. #Video recap: https://www.youtube.com/watch?v=dDwljQWGe1w&amp;feature=youtu.be More photos: https://benjaminhan...
-
Post #2839061
“‘At some point between #marathon and #ultramarathon distances, the damage really starts to take hold,’ said Travis Nemkov, … ‘We’ve observed this damage happening, but we don’t know how long it takes for the body to repair that damage, if that damage has a long-term impact and whether that impact is good or bad.” Does running long distances prematurely age you? This study has the answers https://www.runnersworld.com/uk/news/a70462145/ultramarathon-red-blood-cell-study/ #running #health
-
Post #2839060
“The order mandates agencies to explore a range of policy options, including severance standards, expanded unemployment insurance, job retraining programs aimed specifically at white-collar workers, worker ownership models and a concept the governor called “universal basic capital…” ” After Meta Layoffs, Newsom Signs AI Order to ‘Protect Workers’ and Jobs | KQED https://www.kqed.org/news/12084655/after-meta-layoffs-newsom-signs-ai-order-to-protect-workers-and-jobs #jobs #ai #futureofwork #soci...
-
Post #2400612
AVeriTeC (NeurIPS 2023): 4,568 real-world fact-checked claims, web-retrieved evidence, four-way labels, temporal-leak-free split. Two structural gaps: gold answers are frozen but the retrieval surface isn&#39;t (two systems a year apart hit different Google), and the not-enough-evidence class rewards weak retrievers — predicting NEI when retrieval fails matches gold by coincidence. https://benjaminhan.net/posts/20260507-averitec/?utm_source=mastodon&amp;utm_medium=social #Paper #Bench...
-
Post #2400611
Could #agentic #coding be a &quot;self-undermining loop&quot;, a version of what Anthropic calls the paradox of supervision? To supervise it you need the exact skill it erodes. @tao has made a parallel point for #math: #AI is racing through generation and verification while the human&#39;s digestion lags. Possible solution: Decision-bound Programming. Tools should narrow the space the human must evaluate, not maximize generated code. https://benjaminhan.net/posts/20260507-agentic-co...
-
Post #2400610
I’m so proud that I just decreased the rendering/deploy time of my #blog https://benjaminhan.net by 50%! #softwareEngineering
-
Post #2400609
Jack Clark puts 60% on fully automated AI R&amp;D by end of 2028, 30% by 2027. The case: benchmarks for every sub-skill trending up — coding (SWE-Bench ~2% → 93.9%), training-loop optimization (2.9x → 52x speedup, human 4x baseline passed three generations back), #METR time horizons (~30s in 2022 to ~12h today). The 30-vs-60 gap is a bet on how often a year-scale human insight still cracks a paradigm. https://benjaminhan.net/posts/20260508-import-ai-455-automating-ai-research/?utm_source=ma...