🏆 "Data Drums" won a 2026 Data Sonification Award! I'm excited to continue exploring the overlaps between percussion performance and data storytelling 🪘👏🏽📊 Big thanks to my amazing co-creators Marcus Santos & Lily Gabaree and the performers. https://dataculture.northeastern.edu/2026/04/29/data-sonification-award.html
Creative data & computational journalism. Community Data" book out! https://communitydatabook.com
Posts
Another great venue for bringing together journalists, academics & industry is ISOJ. Coming this Sept 17 & 18 in Austin. Registration is open now https://knightcenter.utexas.edu
Work in local journalism? Later this month my colleague Dan Kennedy is again bringing together innovators for his What Works: The Future of Local News event. Free virtual skill-building webinar on May 21 https://www.eventbrite.com/e/audience-ai-and-events-tickets-1987362369345
"Software Brain": this is one of the best pieces I've read about the current idea of genAI in the world. Patel connects back to critiques I remember from the Big Data era. This framing is already help me diagnose points of view as I move between camps of AI objectors & enthusiasts. https://www.theverge.com/podcast/917029/software-brain-ai-backlash-databases-automation
Related to my earlier post about metrics and questioning if bigger models performs better: here's a nice write up from Rest of World on "Frugal AI" approaches outside of western contexts. https://restofworld.org/2026/frugal-ai-big-tech
📆 Next virtual data crafting circle from DVS is Monday night the 27th. Sadly it conflicts with the in-person studio nights we run in our garage, so I can't join. But maybe you're free? https://www.eventbrite.com/e/data-craft-circle-tickets-1984699394315
More Media Cloud behind-the-scenes: this new viz shows us how stories move through our ingestion system. About 85% of the URLs we fetched over two days made it into our index. Great work from our dev team 👏🏽 I love a good sankey diagram!
It was my birthday recently and I thought all you data nerds might enjoy this card my family got me last year 🤣
Project upgrade: explore a partisan breakdown of top words used in American online news headlines each week. Using Media Cloud data. https://dataculture.northeastern.edu/our-own-terms/
Some initial things I noticed:
* "illegal" is consistently used on by the right in headlines
* "trump" still dominates headlines; the right even writes about him over xmas(!)
* "American" is used more in headlines on right
* and a sanity check: football coverage is seasonal
Full data available at https://github.com/dataculturegroup/on-our-own-terms
Recently upgraded my in-office Media Cloud internals data dashboard display to an even bigger one! Now I gotta get a proper wall mount for it so I'm not worried it will slip and fall 🤞🏽 Powered by a tiny Raspberry PI. Shows how our data pipeline and servers are performing live.
My spring Data Culture Group newsletter update just went out! Subscribe now to get occasional updates on my research in computational journalism, community data, digital humanities, public interest tech, and more. https://dataculture.beehiiv.com/
Curious about the latest learnings, activities & results from our data theatre work? Join in person or online for a session on "Civic Data Theatre and Municipal Climate Action" on April 14th. https://camd.northeastern.edu/events/civic-data-theatre-and-municipal-climate-action/
At a local TV newsroom? My colleague Mike Beaudet is recruiting another round of partners to host innovation fellows working on new visual forms for younger audiences. Watch the trailer to learn more. https://www.youtube.com/watch?si=nHpZA-YyNJyA4GPQ
Apply at https://docs.google.com/forms/d/e/1FAIpQLSdZ4FK1hDQcjnSd462kPD2yoRqeIysJe_T4Zqac77CILQICbQ/viewform
Weekend read: We're rethinking what "accuracy" means for automated classifiers in an AI era. Bigger isn't always better. Let communities define success. Evaluate systems, not algorithms. Work with Yujia Gao on the Counterdata Network with Pregnancy Justice.
https://medium.com/data-culture-group/beyond-accuracy-a-case-study-of-sociotechnical-ml-evaluation-199a48bce664
Media sources change publication volume regularly, so identifying "healthy" sources in Media Cloud is a constant challenge. Paige on our tech team has been exploring if the PELT algorithm can help classify story output over time to catch cases where we lose story feeds from a source. Here's an example from our "most visited US news sites" collection.
If I didn't reply to your emails yesterday, here's why... I was leading the Grooversity Brazilian drum troop at the 🇧🇷 vs. 🇫🇷 friendly, helping Nike's jumpman23 brand launch the new "Joga Sinistro" Brazilian kits. Amazing crowd energy. And now you can see my "drummer concentrating" face in print 🤣 Shoutout to Marcus Santos, JoBeth, Cameron and the other amazing people that made the show happen.
Weekend read: amazing piece in Storybench on trust and dataviz. Namira weaves in perspectives from Cairo, Tufte, and Diehm:
Tufte: Focus entirely on content.
Cairo: Build systematically over time.
Diehm: Create space for conversation.
https://www.storybench.org/how-designers-decide-what-makes-data-stories-feel-trustworthy-before-youve-even-read-the-headline/
More Media Cloud Directory maintenance this week: Phil on our tech team found and removed almost 500,000 "ghost" sources 👻! These were leftovers from our legacy system that had no stories, no feeds, & were in no collections. This should make finding the sources you are looking for easier 👍🏽
Local data journalism funding opp: Data-Driven Reporting Project (DDRP) next deadline is March 31. "document-based investigative projects that serve local and/or underrepresented communities in the US and Canada" https://datadrivenreporting.medill.northwestern.edu