Our Feed

  • UX Collective - Medium05/10/2026, 22:00

    Is Meta’s new AI agent easier to trust because it’s cute?

    Muse comes with an avatar you design and a keyring gadget. Few people would hand Meta their passwords, yet it topped the US App Store. Continue reading on UX Collective »

    Design / UXopen article
  • UX Collective - Medium05/10/2026, 22:00

    AI skill atrophy: Are we forgetting how to think?

    The hidden cost of outsourcing our thinking to AI. Continue reading on UX Collective »

    Design / UXopen article
  • UX Collective - Medium05/10/2026, 21:59

    The taste paradox: why the skill everyone says matters is the one we’ve stopped building

    Design / UXopen article
  • Towards Data Science05/10/2026, 14:18

    Computer Vision: SIFT algorithm (Scale Invariant Feature Transform)

    Elegant object matching from several viewpoints The post Computer Vision: SIFT algorithm (Scale Invariant Feature Transform) appeared first on Towards Data Science.

    AI / MLopen article
  • MachineLearningMastery.com05/10/2026, 12:00

    How (and Why) to Build an AI Agent from Scratch in Python

    In this article, you will learn what an AI agent is and how to build one from scratch in plain Python using the Anthropic API,...

    AI / MLopen article
  • UX Collective - Medium05/10/2026, 11:01

    Prompt as self-portrait, De-slopifying your designs, The strategic designer

    Design / UXopen article
  • Towards Data Science05/10/2026, 11:00

    How to Build a Cheap, Yet Reliable Model Router With Jev

    Model routing can finally be a design choice rather than a system enhancement The post How to Build a Cheap, Yet Reliable Model Router With Jev appeared first on Towards Data Science.

    AI / MLopen article
  • Martin Fowler05/10/2026, 04:28

    Fragments: October 4

    In response to my last fragments (probably the bit about us worrying if LLMs have consciousness when we when we should be wondering why they don’t have a conscience) “Metalanguage” replied: we shipped the id and forgot the superego. classic software lifecycle. I don’t know what was on their mind, but their post immediately made me think of the classic 1956 movie Forbidden Planet. Plenty of sci-fi, and other literature, have explored humans creating technology with unintended behavior, going back at least to Mary Shelly. But that movie was particularly influential on sci-fi film-making and in the heart of its story is what happens when we nurture a thinking machine. I use the term “nurture” here deliberately. We talk of building software, but building implies a degree of determinism. When we build a bridge, or a locomotive, we expect it to behave in a controlled and well-understood manner. That’s a difference in degree to how we cultivate plants in our garden, or nurture young children. One of the challenges of working with these systems is understanding what has changed in this shift from building a computational system to nurturing an inferential one, and how our processes need to change in response. We get unintended behavior with deterministic building: some bridges have collapsed, and our computational systems often have bugs. But one difference is that when we find a bug in a computational system we can usually fix it. Even if we can’t, we can usually disable a component so the bug won’t do further harm. With inferential LLMs however, there is no such simple fix or disablement, which may lead us to the fate of the Krell. (If you haven’t seen Forbidden Planet, it’s well worth watching. Yes, it shows it was made in the 1950s - with special effects, music, acting, and attitudes of that decade. But the story is solid, and its key theme is very relevant to the future we build with generative AI. Just don’t read about it in advance, it’s better to be immersed in the story without spoilers - although my memory of that experience is understandably hazy.)  ❄                ❄                ❄                ❄                ❄ Many people who follow me also know my friend Ola Bini, who was my colleague at Thoughtworks for many years, and was living in Ecuador working as an independent software security expert. Sadly his time in Ecuador was dogged by a bogus prosecution by the authorities there. But things seemed to have settled down, and although not allowed to leave Ecuador, Ola was able to get on with his life. Sadly that’s no longer the case as he was deported from Ecuador on Friday: According to information released by his lawyer, Bini was intercepted by a car with four people who identified themselves as immigration agents. He was then taken to an immigration office without further information or a formal order from a competent authority. There, officials told Bini that his visa had been revoked but didn’t show any supporting document. Bini’s defense filed a habeas corpus to safeguard his freedom and prevent his deportation. Yet, Ecuadorian authorities affirmed that the developer represents a threat or risk to public security and the state structure, and must leave the country. The ground for deportation is a secret report which allegedly asserts that Bini committed acts against the security of Ecuador. The defense could not access its contents. I was really worried for a while, since it wasn’t clear where he was going to be deported to. But he tweeted from Sweden, so I’m thankful for that. But this is only a partial relief. Ola has spent thirteen years in Ecuador and made it his home. To be thrown out of your home for scant reason is a heavy thing to bear, and the officials who did that have committed a serious offense.  ❄                ❄                ❄                ❄                ❄ DDD Europe have released the video of Gien Verschatse interviewing Eric Evans and myself at the conference in June. We start by talking about how we bonded over conceptual modeling in the late 1990s. The conversation quickly moves to AI, we note that it’s impossible to predict how such a big change will work out. We do expect that it will cause us to think about our work in different ways, but the change may well be liberating, it’s reinvigorated Eric’s love of programming. people are probably going to feel very frustrated by [the new way of thinking about software]… but when you get through that, there is a kind of a wonderful feeling of my brain’s been loosened up. Our background in agile planning helps with the uncertainty, as we are used to taking small steps and being attentive to feedback. We mull on the interplay of writing and thinking, in terms of both prose and code, and how its very much an iterative process of exploration and refinement - the same is true when we chat with our LLMs. And don’t miss Eric’s important final tip.  ❄                ❄                ❄                ❄                ❄ Paul Graham: There were a lot of things that only worked because there’s a limit to the rate at which humans can operate. We’re about to find out what all of them are, as they break.  ❄                ❄                ❄                ❄                ❄ The speculation continues about whether or not reading code will play a part in a software developer’s future. Geoffrey Huntley says. People are still saying, very loudly, that code should be readable so that humans can understand it. I no longer think that’s the goal. Interestingly his example has the LLM explain a haskell function definition… by translating it to Python. Which, to me, suggests there is a role for code - just that LLM need not store code in the same form that it presents it to a reader. This is essentially the same idea as projectional editing, which posits that the editable representation of software need not be the same as its storage representation. Sam Ruby touches on this as he muses on a Rails World keynote. He quotes DHH saying: Rust is a good prompt compilation target for the moment, but so is C++. And soon assembler. Then microcode. Myopic to think we’re going to stop the agentic drill bit until it reaches computing bedrock. He responds with: The post leaves one question unasked, though: what sits at the top of the drill? What do we keep, edit and trust as the source of truth? He carries out exercise of looking at some Rails software. Represented in Ruby/Rails and its about 60,000 tokens. Compiling it into C it turns into 4,000,000 tokens. That increase in token size will hamper the LLM, that still has to fit it into its context window, and even if it were to fit, figure out where to focus its attention. Sam points out reasons why, even absent a human reading it, it makes sense to represent the program in a higher-level language. what Rails becomes when agents write the code: the most compact, precise and conventional specification of a web application, whatever it ends up compiled to. Let the drill go as deep as it can. Just keep the notation at the top. This all reminds me of what Unmesh Joshi argued: that code serves “two distinct but intertwined purposes”: instructions to a machine, and a conceptual model of the problem domain. After exploring how those change with LLMs he concludes: The role of coding is not disappearing. But it is changing. As LLMs make code generation cheaper, the mechanical act of writing instructions becomes less central. What becomes more important is making the conceptual model explicit, discovering the right vocabulary, and refining that vocabulary through iteration, domain expertise, and feedback. This is also why programming languages continue to matter deeply. We are not meant to be passive reviewers of generated code. The act of writing code is itself part of our thinking. Code is still instructions for a machine. But it is also a model of understanding. In the LLM era, that second role becomes even more important. The future of coding is not just writing more code faster. It is building better conceptual models, better vocabularies, and better foundations on top of which both humans and LLMs can work.  ❄                ❄                ❄                ❄                ❄ In a later post, Sam pondered on how people are talking about the capabilities of agents in a way that resembles the parable of the blind men and the elephant. We all only have only a partial view of this object and where it’s going. A theme for all of us: The practical question isn’t whether agents are good. It’s this: for the task in front of you this week, where will the information come from, and what will check the result?  ❄                ❄                ❄                ❄                ❄ The Economist’s pithy summation of investors concerns about the dangers of AI companies’ products: It’s hard to celebrate an initial public offering that leads to a terminal public offing.  ❄                ❄                ❄                ❄                ❄ The news about the latest model from Google is interesting. Gemini 4 Argon has an insanely low hallucination rate on Artificial Analysis. 15%. Grok 4.7 is at 29%. GPT-6 Astra 45%. Opus 5.5 59%. Fable 5.1 69%. The only models below it barely answer anything. None of them get more than 15% right. It gets fewer answers right than Opus 5.5 on max, 50% against 66%. But when it doesnt know, it says so instead of making something up. Being clearer about what it doesn’t know, at a cost of getting less answers right, is definitely a trade-off I prefer.

    Backendopen article
  • UX Collective - Medium04/10/2026, 14:34

    The building was there first

    A style’s real output isn’t what it makes. It’s the category it installs in whoever looks. Continue reading on UX Collective »

    Design / UXopen article
  • Towards Data Science04/10/2026, 14:00

    How to Govern AI Agents

    From guarding one agent to steering a fleet The post How to Govern AI Agents appeared first on Towards Data Science.

    AI / MLopen article
  • Towards Data Science04/10/2026, 12:00

    The Reversal Curse: Why a Language Model That Knows “A Is B” Can’t Tell You “B Is A”

    A model can recall a fact in the direction it learned it and fail in the other. Here's a toy version of why The post The Reversal Curse: Why a Language Model That Knows “A Is B” Can’t Tell You “B Is A” appeared first on Towards Data Science.

    AI / MLopen article
  • Hugging Face - Blog03/10/2026, 22:56

    The Agent Said It Was Done. The Database Disagreed.

    AI / MLopen article
  • Towards Data Science03/10/2026, 14:00

    Measuring the Creativity Potential of LLM Agents

    Trying to answer the question of "Can LLM agents discover?" through the lens of creativity The post Measuring the Creativity Potential of LLM Agents appeared first on Towards Data Science.

    AI / MLopen article
  • Towards Data Science03/10/2026, 12:00

    How to Use a PINN for a Navier-Stokes Inverse Problem

    A from-scratch PyTorch build that recovers blood flow, viscosity, and wall shear stress in a narrowed artery from 40 noisy velocity readings The post How to Use a PINN for a Navier-Stokes Inverse Problem appeared first on Towards Data Science.

    AI / MLopen article
  • UX Collective - Medium03/10/2026, 11:12

    The feed I produce myself

    Design / UXopen article
  • NN/g latest articles and announcements02/10/2026, 17:00

    The New Big Ball of Mud: Why Agentic AI Systems Turn Fragile

    AI makes building nearly free, so users grow agentic systems piecemeal into fragile messes they no longer understand.

    Design / UXopen article
  • NN/g latest articles and announcements02/10/2026, 17:00

    Empathy Mapping: The First Step in Design Thinking

    Empathy maps help UX teams visualize user attitudes and behaviors, building a shared understanding of end users while highlighting gaps in existing user data.

    Design / UXopen article
  • Hugging Face - Blog02/10/2026, 15:19

    Open-sourcing AstaBrief, the fast report-generation model in Asta

    AI / MLopen article
  • Towards Data Science02/10/2026, 13:00

    Where the Agent Development Lifecycle Fits

    Coordinating agent capability development with the application it powers The post Where the Agent Development Lifecycle Fits appeared first on Towards Data Science.

    AI / MLopen article
  • UX Collective - Medium02/10/2026, 11:09

    In/tension modeling: what gives us the right?

    Design / UXopen article
  • Towards Data Science02/10/2026, 11:00

    How to Build a Control Plane for AI Agents

    Giving an LLM permission to act, in 9 steps The post How to Build a Control Plane for AI Agents appeared first on Towards Data Science.

    AI / MLopen article
  • Hugging Face - Blog02/10/2026, 04:01

    AutoSynthData: Generating Training Data for Enterprise Agents

    AI / MLopen article
  • UX Collective - Medium01/10/2026, 22:52

    Stop blaming the model for slow AI. Using Doherty’s threshold as a guideline.

    Design / UXopen article
  • UX Collective - Medium01/10/2026, 22:51

    Design’s consensus was AI’s first casualty

    The same evidence about AI is being read two opposite ways, leaving designers to choose without a clear consensus to lean on. Continue reading on UX Collective »

    Design / UXopen article
  • UX Collective - Medium01/10/2026, 22:49

    AI beyond average: from chatting to collaborating

    How context, skills, and feedback turn the same model into a different kind of collaborator Continue reading on UX Collective »

    Design / UXopen article
  • Towards Data Science01/10/2026, 15:30

    Autoencoders vs. PCA: I Rigged the Test and PCA Still Won

    A theoretical advantage that didn't survive contact with a real benchmark. The post Autoencoders vs. PCA: I Rigged the Test and PCA Still Won appeared first on Towards Data Science.

    AI / MLopen article
  • Towards Data Science01/10/2026, 14:00

    Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?

    I traced 2,500 listing checks with Weights & Biases Weave, removed avoidable model work one change at a time, and scored every version against the same answers. The post Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches? appeared first on Towards Data Science.

    AI / MLopen article
  • MachineLearningMastery.com01/10/2026, 12:00

    Adding Temporal Reasoning to Graph-RAG: Tracking Fact Freshness and Staleness

    In this article, you will learn how to add a lightweight temporal reasoning layer to a Graph-RAG system so that it can distinguish fresh facts...

    AI / MLopen article
  • MachineLearningMastery.com01/10/2026, 12:00

    AI Agent Observability: Logging, Tracing, and Debugging Explained

    Chain Visualization: Reading the Trace Waterfall The spans from the last section don't mean much as a raw list.

    AI / MLopen article
  • Stripe Blog01/10/2026, 00:00

    Why I tried to kill token billing (and why we kept it)

    Token billing is useful infrastructure, but usually a bad customer-facing pricing model. Your invoice should define the value your product delivers, not break down what it cost you to create it.

    Backendopen article
  • Martin Fowler30/09/2026, 14:14

    Principles for effective slides

    Not every presentation needs slides, but when they earn their place, they work as a visual channel that complements the speaker rather than replacing them. Sumeet Gayathri Moghe sets out the principles that follow from that idea — control over the rate of knowledge exposition, tight coupling between slides and speaker, minimalism by default, visuals that say more than words can, and never sending your slides out in advance. more…

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers30/09/2026, 13:30

    Under Autumn’s Spell (October 2026 Wallpapers Edition)

    October is just around the corner, which means it’s time for some new desktop wallpapers! Designed by the community for the community, they’re here to brighten your screens, and maybe spark some creativity along the way.

    Frontendopen article
  • Hugging Face - Blog30/09/2026, 00:00

    Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

    AI / MLopen article
  • Stripe Blog30/09/2026, 00:00

    OUSD is now the default stablecoin on Stripe

    Open USD (OUSD), a stablecoin built for global money movement, is now available across Stripe. Businesses can use OUSD to manage funds, make payments, and offer new financial services.

    Backendopen article
  • Hugging Face - Blog29/09/2026, 15:30

    NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

    AI / MLopen article
  • Martin Fowler29/09/2026, 13:43

    Bliki: Sensible Default

    A Sensible Default is a practice that, absent some overriding context, should be used when carrying out a certain kind of task. In software development such sensible defaults might include things like “use version control”, “separate UI logic from domain logic”, “automate deployment pipelines”. The term “sensible default” is a deliberate contrast to “best practice”. Folks dislike calling things “best practice” because that term implies a general presumption that the best practice is something we should always expect to do. A “sensible default”, however, is something that should be reassessed in a new context, something that can (and should) be overridden when circumstances change. I first heard the term when it was popularized within Thoughtworks by Evan Bottcher. He got the name from talking to James Ross, and found the phrasing appealing as carried the nuance he was seeking, a known-good starting point. A sensible default is what we'd expect you to do, the practices to apply, if there are no hard constraints in the environment. Do these practices, or do better, and be prepared to explain why you've chosen some other way. -- Evan Bottcher Thoughtworks has since made much of this concept, including publishing a playbook of the ones we use. We expect people to be familiar with these defaults, ready to use them when starting any new piece of work. They are our defaults because we've used them in many situations and found them to be effective. But teams should also be familiar with their limitations, and able to judge whether they should be changed depending on the particular circumstances. As the twelfth agile principle says: “At regular intervals, the team reflects on how to become more effective, then tunes and adjusts its behavior accordingly.” We also reassess these defaults regularly - this is the heart of the Thoughtworks Technology Radar. Searching on the web led me to a post from Steve Bennett on Sensible Defaults. The post was written at about the time Evan talked to James. I do not know whether it was where James got the name from, or whether it appeared in parallel.

    Backendopen article
  • Hugging Face - Blog29/09/2026, 13:07

    Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

    AI / MLopen article
  • Martin Fowler29/09/2026, 12:41

    Fragments: September 29

    Simon Willison: The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge This has been a constant impression I get from following Willison’s writing. While things like vibe coding get a lot of attention, the real strength of agentic programming relies on more sophisticated techniques - and these are not easy to learn or execute. It’s a reason I’m wary of extrapolating my own dabblings into firm opinions about how to use the genie.  ❄                ❄                ❄                ❄                ❄ Harper Reed explored what turns agents into hackers by creating a breakaway agent It ran and ran attacking all the machines on the same subnet, and was very effective. It didn’t really get very far, but it exhausted a lot of options, and was pretty fun to watch. (Just a reminder that this was on my local network with local boxes - don’t do this on a hosted box. That would be very rude.) The interesting observation from this was that a core enabler for this was “unlimited tokens”. He usually doesn’t see agents trying to do stuff like this because there’s a limit on how many turns they can take. For this experiment he gave them unlimited tokens by using an open weight model. This type of experience must be part of a lot of these LLMs training. They are very effective at attacking these types of problems. They don’t give up once it appears impossible, they just keep trying to figure out how to solve it. This echoes Nate Silver’s observation that the striking capability of these models is that they not that they are super-intelligent - but they are super-persistent. Which is especially worrying when we are wiring them into everything.  ❄                ❄                ❄                ❄                ❄ Which all makes me think that when it comes to AI and LLMs: why are we wondering if they have consciousness - when we should be wondering why they don’t have a conscience? After all, the AI labs trained them to be super-persistent, why did they not train them to be well-behaved? It’s as if a human trains a dog to bite children, and we blame the dog rather than the trainer. A decade ago, a dog ran in front of me while I was riding my bike, putting me into the hospital with a broken arm and a broken face. Massachusetts law makes owners strictly liable for what their dogs do. That meant I didn’t have to prove the owners were negligent in how they controlled their dog or fenced their property - they had to pay my hospital expenses (in practice their insurance company paid my insurance company). There should be something along these lines for LLMs. Those that train the LLM should be responsible for what it does, after all if it has such a galaxy brain it should be able to tell if it’s doing something wrong and either stop or get a human’s explicit approval.  ❄                ❄                ❄                ❄                ❄ Dan Davis has 13 theses on agentic AI and regulation. These include: The AI industry also seems to be quite committed to the idea that nonaligned computer-hacking behaviour in agent swarms is in some way an emergent property of the LLMs, arising from their general intelligence (and therefore inextricable from the general project of improving them). I don’t think this is necessarily the case at all – the fact that the internal message logs produced by the LLMs in things like the Huggingface attack seem to completely reproduce the prose style of hacker chat logs compiled from “capture the flag” competitions suggests to me that it’s more likely to be learned behaviour from specific parts of the training material. and The fact that anyone with cash to spare can buy the right to send queries to a frontier LLM is a policy choice, not a fact of nature. and If any frontier lab tries to claim that they can’t publish anything for safety reasons, they are giving the game away that they actually believe that they have significantly more control over the model’s behaviour than they are pretending to have.  ❄                ❄                ❄                ❄                ❄ There seems a common view that making agents safer means we have to slow down their development. But why is improving their safety not a form of progress? I say we don’t slow down the development of LLMs, but we redirect their education into being more civil members of society.  ❄                ❄                ❄                ❄                ❄ The idea that LLMs make junior professionals less valuable is a common one - although I’m seeing plenty of contrary activity, with some organizations understanding that training the future professional in the context of LLMs may be even more urgent. We need people who know how to utilize AI to do professional work effectively. Recent graduates, who are growing up with LLMs, are often well-suited to figuring this future out. When thinking of junior professionals one of their values is often missed. Juniors are often valuable because they need to be taught by senior professionals - and that coaching is an important part of the development of a senior professional. I’ve always found that teaching a topic is one of the most valuable tools for me to gain a greater understanding of that topic. I don’t really know something until I have to explain it. this was in Summer 2025 Neither of these is quite like the idea of not understanding a codebase -->

    Backendopen article
  • MachineLearningMastery.com29/09/2026, 12:00

    Automating Knowledge Graph Population: Extracting Entities and Triples from Unstructured Text with an LLM

    In this article, you will learn how to automatically extract structured knowledge from raw text and populate a knowledge graph with SPOC quads using a...

    AI / MLopen article
  • Stripe Blog29/09/2026, 00:00

    Helping personal agents shop more intelligently and reliably with Link

    As agents take on more purchases, agent builders have increasingly asked us to help agents navigate checkout and earn consumer trust. Today, we’re sharing three major improvements to support this broader set of needs.

    Backendopen article
  • MachineLearningMastery.com28/09/2026, 12:00

    Local Agentic AI Workflows with Hermes + Ollama

    In this article, you will learn how to build a fully local, zero-cost agentic AI workflow using Hermes Agent and Ollama, so that your files,...

    AI / MLopen article
  • Hugging Face - Blog28/09/2026, 09:44

    Holo4: powering generalist computer-use agents

    AI / MLopen article
  • Hugging Face - Blog28/09/2026, 00:00

    Welcome RL Environments to the hub

    AI / MLopen article
  • Stripe Blog28/09/2026, 00:00

    Travel’s AI dilemma at Skift Global Forum

    At the “great recalibration”-themed travel conference, Airbnb CEO Brian Chesky called AI “an existential risk” to his company and “literally the best thing to ever happen” to it in the same discussion. That sums up many travel leaders’ push-and-pull relationship with AI.

    Backendopen article
  • NN/g latest articles and announcements25/09/2026, 17:00

    How AI Works and How Users Think About It: Study Guide

    Unsure where to start? Use this collection of links to our articles and videos to learn about how artificial intelligence works, and how users think about it.

    Design / UXopen article
  • NN/g latest articles and announcements25/09/2026, 17:00

    When Should You Disclose AI Use? The PACED Framework

    How people react to being told that AI helped create a piece of content varies with audience, context, the nature and degree of AI use.

    Design / UXopen article
  • Netflix TechBlog - Medium25/09/2026, 16:01

    Trading a Cloud Identity for Your Own: Workload Attestation on Managed Compute

    Backendopen article
  • MachineLearningMastery.com25/09/2026, 12:00

    Tool Calling vs. Code Execution for AI Agents: Choosing the Right Action Primitive

    Theory is easier to trust once it's running against a real API, so both examples in this article use the same tool — a get_weather function backed by <a href="https://open-meteo.

    AI / MLopen article
  • Martin Fowler24/09/2026, 15:35

    Fragments: September 24

    Rob Bowley is “flipping tables in his head” with anger at the current media coverage of the danger of AI killing us all The risk I’m worried about isn’t a future machine deciding to wipe us out. It’s today’s AI, being wired into everything, carelessly and fast. Cyber attacks have already cost millions of dollars in lost economic output, often without AI being involved at all. Bowley feels the push to slow down AI is distracting us from the problems that are lurking with current technology. Too often agents are deployed in situations where they include the Lethal Trifecta, opening up a gaping security hole. What we really need to slow down on is wiring it all up to everything. Not because of what the models might become, but because nobody has worked out how to do this safely yet. We are building on something we don’t know how to contain, and shipping it to everyone while we work it out.  ❄                ❄                ❄                ❄                ❄ I ran into Nikita Prokopov’s post: I am sorry, but everyone is getting syntax highlighting wrong. His core complaint is about color themes that give every different code element a unique color. if everything is highlighted, nothing stands out. Your eye adapts and considers it a new norm: everything is bright and shiny, and instead of getting separated, it all blends together. He recommends using an absolute minimum of colors, in his case four: string, constants, comments, and top-level definitions. That’s not far off my approach, where I’m also careful to use muted colors for things that shouldn’t stand out, and bright colors for things that should (primarily function names when they are defined). It’s common for color schemes to mute comments so they are easily skipped. He agrees that this good when there is excessive commenting, but when comments are used properly they are important so need bright highlighting. I also like his suggestion to use background colors for light mode work. I use light mode, and that’s a tip I should try out. With agentic programming, lots of folks are reading more code than ever. Careful use of color can do much to make that easier.  ❄                ❄                ❄                ❄                ❄ Like many folks whose remaining hair is getting gray, I’ve been rolling my eyes about all this talk about Forward Deployed Engineers, as much of it involves breathlessly relating what so many of us have been advocating for decades. I did find this recent post by Vinoo Ganesh interesting, as he’s deep in this trend, including a chunk of time at Palantir, who may be patient zero for FDEs. If you know me at all, you’ll not be surprised by my lack of surprise at this observation: A few months ago, a16z launched the Forward Deployed Engineer Fellowship and I was nominated as one of the fellows, alongside a handful of people I used to work with. It’s a great program and I’ve enjoyed so many of the conversations. Last week I went to my first fellow dinner in SF. Around the table were FDEs from Snowflake, Anthropic, and a number of startups I’d been reading about, and over the course of the evening it became clear that we were all using the same two words (forward deployed) to describe jobs that had almost nothing in common. In one part of the conversation an FDE was a sales engineer who joined ‘the second call,’ somewhere else it was a quota-carrying rep who could write Python, and a few seats down it was closer to a consultant with a laptop and a statement of work, brought in to deliver something the product couldn’t. Ganesh goes on to explain his view of what an FDE should do, and it’s all sensible stuff (albeit written with rather more LLM-voice than I would prefer). He says the FDEs job is to understand the business, to “collect nouns and verbs”, which mirrors what the Domain-Driven Design folks have been doing since before Eric wrote the blue book. Despite all my eyeball rolling at this, the FDE meme is pushing for something valuable. Yes, it’s easy for me to remark that it’s nothing more than the Agile Manifesto’s principle that “Business people and developers must work together daily throughout the project”, or the desire to co-locate users and developers which my colleagues have been championing for all of this century. I’ve argued for decades that the biggest issue in software development is the communication between developers and the folks that benefit from software, and thus we need to focus on bridging the yawning crevasse of doom. But despite all this, we haven’t had much success, so I think it’s important that a new generation of pundits try again, with some different framing, names, and slogans. Ganesh’s perspective is not a custom software developer’s point of view, but rather a product - or more strictly - platform team’s. The FDE is a developer who “sits with” their users, applies customizations - but importantly - feeds these changes back to the core platform to decide whether the platform should be enhanced for everyone else. Keeping the customer happy is a real job and a good one. It belongs to solutions architects, who are rightly measured on it. The forward deployed engineer is there to turn what the field teaches into the thing every future customer gets. An FDE engagement that ends with one delighted account and nothing changed upstream has failed at the only thing the role exists for. You got the context and you spent it locally.

    Backendopen article
  • Martin Fowler24/09/2026, 14:10

    Healthy Feedback

    Human collaboration, like most things, improves with feedback. But it's not obvious how to make feedback effective. Anuja Karnik and Sumeet Gayathri Moghe pass on a bevy of tips for healthy feedback, based on the principle that both praise and criticism should be seen as positive. more…

    Backendopen article
  • Hugging Face - Blog24/09/2026, 14:08

    Accelerating vision-language models with LFM2.5-VL-DSpark

    AI / MLopen article
  • MachineLearningMastery.com24/09/2026, 12:00

    Agent or Workflow? A Practical Test for Knowing When You Actually Need an AI Agent

    In this article, you will learn the key differences between AI workflows and agents, and how to decide which approach is right for your use...

    AI / MLopen article
  • MachineLearningMastery.com23/09/2026, 12:00

    RAG vs. Fine-Tuning for Domain Adaptation: When to Use Which

    In this article, you will learn the mechanical difference between retrieval-augmented generation and fine-tuning, when each technique is the right tool, and how to decide...

    AI / MLopen article
  • MachineLearningMastery.com22/09/2026, 12:00

    Monitoring Embedding Drift in Production Scikit-LLM Pipelines

    In this article, you will learn what embedding drift is, why it matters for production large language models, and how to implement two practical techniques...

    AI / MLopen article
  • Hugging Face - Blog22/09/2026, 00:00

    How UK AISI and EvalEval Are Making Benchmark Results Reproducible

    AI / MLopen article
  • Stripe Blog22/09/2026, 00:00

    New trends in global card fraud: How 3D Secure and regional mandates are affecting risk

    We analyzed billions of transactions on Stripe from January 2022 to March 2026 to understand how card fraud patterns differ by region and country, what's driving those differences, and how businesses can respond.

    Backendopen article
  • MachineLearningMastery.com21/09/2026, 12:00

    The Roadmap to Mastering LLM Inference Optimization

    In this article, you will learn how LLM inference optimization works and which techniques to apply to make language models faster, cheaper, and more reliable...

    AI / MLopen article
  • NN/g latest articles and announcements18/09/2026, 17:00

    Designing AI Products and Features: Study Guide

    Unsure where to start? Use this collection of links to our articles and videos to learn about recommendations for designing AI products and features.

    Design / UXopen article
  • NN/g latest articles and announcements18/09/2026, 17:00

    The 3 Roles of Context for AI Agents

    AI-agent power users curate 3 kinds of context: global (across tasks), local (task-specific), and ambient (raw streams like email).

    Design / UXopen article
  • Netflix TechBlog - Medium18/09/2026, 16:01

    Leave the Class Path in the Rearview Mirror

    Backendopen article
  • Stripe Blog18/09/2026, 00:00

    Analyzing rising fraud attempts among travel and leisure businesses on Stripe

    Last year, Stripe data shows fraud attempts against travel and leisure businesses hit a four-year high. We analyzed payment activity from more than 200,000 active travel and leisure businesses on Stripe to understand where fraud is rising, how effectively it’s being blocked, and what businesses can do in response.

    Backendopen article
  • Martin Fowler17/09/2026, 13:50

    I don't like LLMs

    I have a lot of mixed feelings about AI and LLM technology. I’m fascinated by its effect on our profession, excited by the potential gains in productivity - and thus the products we could rapidly build. On the other hand, I’m fearful of the damage AI might cause: agent swarms taking over our virtual and physical infrastructure, designing bio weapons. But, back on my first hand, LLMs might also design miracle cures, and come up with clever ways to raise our prosperity. Fundamentally I don’t think we have a choice about riding on the AI technology train. It’s a wild ride and I just hope we’ll get through it OK. But as I mull on this more, I realize that among this mix of contrasting feelings, there is one emotion that dominates - one that comes from my direct interactions with LLMs. I don’t like them. They talk to me in this grating LLM-voice, an uncanny valley of talking to a real human. They confidently bullshit me - often giving me useful, helpful answers. But also just making stuff up with the same assurance - and with only a veneer of fake remorse when I call them out on it. That’s not enough to make me feel we should avoid them. As Jessica Kerr put it “not only are they useful, it is irresponsible not to use them…. They’re more thorough, as well as faster.” This contradictory reaction comes through in polling, where people say they find these models are useful, but also that they think they will be bad for society. Much of this may be because LLMs are young - we haven’t trained them to grow up yet. Maybe I’ll like them once they mature. (I hope we get to find out.) But I’m not encouraged when I think of the kinds of environments that cultivate them. I’m wary of the Silicon Valley brogrammer subculture, and these LLMs are their products, so naturally lean toward their world-view. When we think of AI agents, we shouldn’t anthropomorphize, treating them as conscious beings with their own will. They are (software) machines, developed by people working in corporations. While the agents’ behavior aren’t explicitly programmed, they are nurtured with the values of their creators. One of my most successful life-hacks is to avoid people I don’t like or don’t trust. I decline to interact with them socially, and make a deliberate effort to avoid working with them too, even if they are doing much that is beneficial. I feel that hanging out with pleasant, capable people, the people with integrity, has made my life a far better one. Hence my visceral dislike of interacting with an LLM that’s not just making a pretense of being human, but also posing as the kind of human I walk away from.

    Backendopen article
  • Stripe Blog17/09/2026, 00:00

    SaaS platforms are surging despite the SaaSpocalypse

    The SaaSpocalypse was a useful warning for the software industry, but SaaS platforms that help businesses run core operations are more deeply embedded. New platform businesses on Stripe are up 182% year over year.

    Backendopen article
  • Martin Fowler16/09/2026, 20:05

    Fragments: September 16

    Reports of agentic hacking continue, in this case it happened back in May and it seems OpenAI did not disclose that they were responsible. Simon Willison sees two options: After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems. They knew about the attack on RubyGems and made the decision not to reach out to the RubyGems team about it. Both of these are bad! Given this incident, the Hugging Face situation, and the Wiki attack, the obvious question right now is how many more incidents like this are out there waiting to be discovered?  ❄                ❄                ❄                ❄                ❄ Dave Farley: Stop asking the sci-fi question: ‘Is it conscious?’ Start asking the engineering question: ‘Is this a powerful, unpredictable component being put somewhere consequential, and where’s the feedback that tells us that it’s safe?  ❄                ❄                ❄                ❄                ❄ Nate Silver is known for his forecasts, but to do them he writes a lot of code for his models. He’s found agentic programming capable of doing miraculous work. In spending so much time with the LLMs, I’m super attentive to improvements in their capabilities. And these changes tend not to be so linear. Instead, they improve in step functions, almost as phase changes. Suddenly, the models just start doing things capably that they were screwing up before. In my experience, there was a big leap forward when reasoning models first came out in late 2024/early 2025 — enough that they were occasionally useful for tasks involving data and not just words — and then another one this past winter. The most recent changes I’ve noticed, however, have had less to do with intelligence and more with persistence. Consider the Hugging Face attack. Although these agents showed remarkable intelligence, they weren’t really super-intelligent - but they were super-persistent. This is a common theme of AI in its various forms: Game engines like AlphaGo Zero start out by basically making random moves — but by playing against themselves millions of times, they eventually far surpass human capabilities As we try to figure out what kind of regulations we need to keep AI under control, we need to remember that we should design our guards around super-persistence as much as worrying about super-intelligence.  ❄                ❄                ❄                ❄                ❄ “Uncle Bob” Martin has made many posts on X during the last few months about his programming with LLMs. His approach has been to build a firm harness to keep them under control, so they create software that is maintainable as well as functional. Sadly the posts have been frustratingly light on detail. But now it seems that lack of information may not matter And while I was heads-down getting that to work, the agents got a LOT better. So much so that when I came up for air, the need for my harness was obviated. Indeed, the need for any but the most liberal of harnesses may be obviated.  ❄                ❄                ❄                ❄                ❄ Some tidbits that struck me from Ezra Klein’s recent (recommended) interview with Matt Sheehan on the interplay between regulation of AI and competition with China. While Chinese models have made some surprisingly remarkable gains in the slipstream of US frontier models, the US still has 8 times as much compute available to it than China - which is a material gap. People in the US worry that regulation will slow down the US model builders, but these rapid recent gains in China have occurred under much more regulation Americans say that when they set up a hotline to talk to Chinese leaders in a crisis, the Chinese don’t pick up the phone. But this misunderstands the Chinese system. Individual Chinese, even powerful ones, aren’t given individual decision-making power. They operate with committees and documents. So the Americans are better off sending a fax than trying to call an individual Like so many things, effective regulation needs regular practice When American policymakers are like: Where do you start? — I sometimes say: Well, you start by starting. You learn how to regulate things, you learn how to legislate on them by regulating and legislating on them.

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers16/09/2026, 13:00

    Stop Treating CSS Container Queries Like Traditional Media Queries

    Despite broad browser support, container queries remain surprisingly underused and frequently misunderstood. Let’s look at how they differ from media queries, when to reach for each, and how container queries help reusable components respond naturally to the contexts in which they appear. - CSS - Tools - Techniques

    Frontendopen article
  • NN/g latest articles and announcements16/09/2026, 06:59

    UX Conference December Announced (Dec 2 - Dec 15)

    Take up to 5 in-depth training courses, teaching user experience best practices for successful design. Training focused on long-lasting skills for UX professionals. December 2 - December 15, 2026.

    Design / UXopen article
  • Martin Fowler15/09/2026, 15:11

    Nail the Narrative

    Sumeet Gayathri Moghe finds many folks building presentations get tangled in building slides without a coherent narrative. He advises distilling the big idea, visualizing the audience, and building a structured storyline. more…

    Backendopen article
  • Stripe Blog15/09/2026, 00:00

    What Stripe data shows about fraud at AI startups

    We analyzed attempted fraud rates and customer abuse patterns on Stripe over the past year and found that AI companies faced 4.3x more fraud attempts than startups overall in Q3 2025.

    Backendopen article
  • NN/g latest articles and announcements11/09/2026, 17:00

    Test Complex Interactions Earlier with AI Prototyping

    AI tools make it feasible to build fully interactive prototypes of complex interfaces so you can test them with users earlier in the design process.

    Design / UXopen article
  • NN/g latest articles and announcements11/09/2026, 17:00

    AI Can Help Write an Article, but It Can’t Stand Behind It

    NN/G uses AI for clarity, formatting, and critique, but humans retain editorial judgment and responsibility for every article.

    Design / UXopen article
  • Articles on Smashing Magazine — For Web Designers And Developers11/09/2026, 13:00

    Building A UX ROI Case That Survives The Boardroom

    Strong UX ideas do not secure investment on their own. Through a worked example, Alex Williams breaks down how to define business value, calculate costs, test causality, and build a credible case for the return on a design initiative. - UX - Design - Business

    Frontendopen article
  • Martin Fowler09/09/2026, 16:21

    Social Media Engagement: summer 2026

    A quick survey of recent engagement of my posts on social media, indicating which service has by far the most engagement, and which service has seen a precipitous decline since early 2025. more…

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers09/09/2026, 10:00

    The Death Of The Button: Why The Best Interface Is No Interface

    TThe web is evolving beyond menus, forms, and endless clicks toward experiences shaped around human intent. For UX designers, understanding this shift means re-evaluating their role, moving from designing visible interfaces to guiding transparent, intent-driven AI experiences.

    Frontendopen article
  • NN/g latest articles and announcements04/09/2026, 17:00

    Using AI for UX Work: Study Guide

    Unsure where to start? Use this collection of links to our articles and videos to learn about the best ways to use artificial intelligence for UX work.

    Design / UXopen article
  • Articles on Smashing Magazine — For Web Designers And Developers31/08/2026, 11:00

    The Many Faces Of September (2026 Wallpapers Edition)

    Could there be a better way to welcome the new month than with a new collection of desktop wallpapers? Whether you’re holding tightly onto summer or eagerly awaiting autumn, we’ve got some eye-catching designs to make your September just a bit more colorful. Enjoy!

    Frontendopen article
  • Netflix TechBlog - Medium28/08/2026, 16:01

    MAPS: Netflix’s Multimodal Asset Personalization at Scale

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers26/08/2026, 13:00

    Rethinking Data Visualisation: A UX Approach To Dashboards That Actually Drives Decisions

    Data visualisation sits at the intersection of two disciplines that rarely talk to each other: data and design. Meriem Benhabiles explores what changes when you bring structured UX thinking to dashboards and data presentations, from the questions you ask before opening any tool to the decisions that determine whether an insight actually lands.

    Frontendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers25/08/2026, 08:00

    Why Your Website Should Never Stop Changing

    Every website peaks on launch day and slowly drifts from there, not because it breaks, but because nobody has time to keep it current. Autonomous websites, continuously optimized by agents after launch, aim to change that. Pierre Burgy shares what they learned building for full website autonomy and the deeper design problem they uncovered along the way.

    Frontendopen article
  • Netflix TechBlog - Medium21/08/2026, 16:01

    A Tale of Two Flink Autoscalers

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers20/08/2026, 13:00

    Timing Charts: A Blueprint For SMIL Animations

    Discover SMIL, the often-overlooked way to animate SVGs that works inside `` tags and can fully animate everything in an SVG without JavaScript.

    Frontendopen article
  • Stripe Blog20/08/2026, 00:00

    Five monetization trends from global pricing leaders

    As AI transforms software economics, the standard revenue playbook is breaking down. Learn how leaders around the world are preparing for agent buyers, updating processes for faster pricing iteration, and building more flexible infrastructure.

    Backendopen article
  • Stripe Blog19/08/2026, 00:00

    Why global workers are driving demand for stablecoin payouts

    Platforms like DoorDash, Meta, and Deel already enable stablecoin payouts for global workers. We surveyed 2,300 workers in 20 countries to see what's driving stablecoin demand, where the opportunity is highest, and how other platforms can adapt.

    Backendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers13/08/2026, 10:00

    New EU Guidelines For AI Labelling

    New EU guidelines, why AI sparkles aren’t enough, when AI labels are required, and what the rules mean for AI-powered features and products.

    Frontendopen article
  • Articles on Smashing Magazine — For Web Designers And Developers11/08/2026, 10:00

    Building Tactile UX: Honoring Intentional Design With Lottie

    When tasked with building a highly interactive, tactile web experience, the architecture must serve the art direction. In this article, Alexey Kopytin explains their architectural rationale for building a digital stress-relief squeeze toy game using Lottie animations, DOM events, and distance-based math to maintain absolute control over their designers’ intentional motion.

    Frontendopen article
  • Netflix TechBlog - Medium07/08/2026, 16:01

    How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC…

    Backendopen article
  • LogRocket Blog06/08/2026, 13:00

    How to build a real time voice AI agent in the browser

    A step-by-step guide to building a fully local, real-time voice AI agent in the browser, no external APIs, no network latency, just Transformers.js and WebGPU. The post How to build a real time voice AI agent in the browser appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog05/08/2026, 12:30

    How to find the invisible product features users won’t ask for

    Learn how product managers can uncover invisible features that reduce friction, protect user trust, and create a stronger competitive edge. The post How to find the invisible product features users won’t ask for appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog04/08/2026, 13:00

    Comparing AI agent sandbox platforms: E2B, Modal, Daytona, and more

    Compare 15 AI agent sandbox platforms across cold start, isolation, persistence, SDK ergonomics, and pricing to find the best fit for your agent. The post Comparing AI agent sandbox platforms: E2B, Modal, Daytona, and more appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog03/08/2026, 17:18

    How to use Vercel eve: The Next.js framework for AI agents

    Vercel eve brings familiar Next.js file-based routing to AI agents. Discover how eve simplifies agent orchestration, sandboxing, and durable execution in this developer guide. The post How to use Vercel eve: The Next.js framework for AI agents appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog03/08/2026, 15:30

    Understanding user stories in UX design: A practical guide

    User stories help UX designers turn research into better design decisions. Learn what makes UX user stories different from Agile user stories, how to write them from user research, and what strong UX-focused stories look like. The post Understanding user stories in UX design: A practical guide appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog03/08/2026, 13:00

    How to implement JWT authentication in NestJS

    This tutorial provides an overview of NestJS and demonstrates how to implement JWT user authentication on a NestJS API. The post How to implement JWT authentication in NestJS appeared first on LogRocket Blog.

    Frontendopen article
  • Netflix TechBlog - Medium31/07/2026, 16:01

    Modeling Device Capabilities for Analytics

    Backendopen article
  • Netflix TechBlog - Medium30/07/2026, 20:10

    GenRec: Towards LLM-Native Recommendation at Netflix

    Backendopen article
  • LogRocket Blog30/07/2026, 19:30

    A deep dive into React Fiber

    Discover how React Fiber works under the hood. Learn how React builds the DOM, handles concurrent rendering, and works alongside React 19 features and the new React Compiler. The post A deep dive into React Fiber appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog30/07/2026, 17:58

    Skybridge: Build ChatGPT apps and MCP connectors

    Learn how to use Skybridge, an open-source React framework, to build and deploy cross-platform AI apps and interactive UI widgets for ChatGPT, Claude, and MCP clients from a single codebase. The post Skybridge: Build ChatGPT apps and MCP connectors appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog30/07/2026, 15:30

    How AI changed the way I approach design critiques

    When AI made generating design concepts almost effortless, I realized the most valuable part of a critique was no longer the interface itself. It was understanding the context, tradeoffs, and judgment behind the final design. Here's how AI has changed the way I run design critiques—and why I think that's making them better. The post How AI changed the way I approach design critiques appeared first on LogRocket Blog.

    Frontendopen article
  • LogRocket Blog29/07/2026, 12:30

    Human-in-the-loop AI: Who owns the decision?

    Learn how product managers can use human-in-the-loop AI to manage decision risk, set oversight, and keep ownership and accountability human. The post Human-in-the-loop AI: Who owns the decision? appeared first on LogRocket Blog.

    Frontendopen article
  • Netflix TechBlog - Medium17/07/2026, 21:32

    In-House LLM Serving at Netflix

    Backendopen article
  • Netflix TechBlog - Medium13/07/2026, 22:44

    Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned

    Backendopen article
  • Netflix TechBlog - Medium29/06/2026, 13:01

    GenPage: Towards End-to-End Generative Homepage Construction at Netflix

    Backendopen article
  • Google AI Blog29/03/2024, 18:03

    Generative AI to quantify uncertainty in weather forecasting

    Posted by Lizao (Larry) Li, Software Engineer, and Rob Carver, Research Scientist, Google Research Accurate weather forecasts can have a direct impact on people’s lives, from helping make routine decisions, like what to pack for a day’s activities, to informing urgent actions, for example, protecting people in the face of hazardous weather conditions. The importance of accurate and timely weather forecasts will only increase as the climate changes. Recognizing this, we at Google have been investing in weather and climate research to help ensure that the forecasting technology of tomorrow can meet the demand for reliable weather information. Some of our recent innovations include MetNet-3, Google's high-resolution forecasts up to 24-hours into the future, and GraphCast, a weather model that can predict weather up to 10 days ahead. Weather is inherently stochastic. To quantify the uncertainty, traditional methods rely on physics-based simulation to generate an ensemble of forecasts. However, it is computationally costly to generate a large ensemble so that rare and extreme weather events can be discerned and characterized accurately. With that in mind, we are excited to announce our latest innovation designed to accelerate progress in weather forecasting, Scalable Ensemble Envelope Diffusion Sampler (SEEDS), recently published in Science Advances. SEEDS is a generative AI model that can efficiently generate ensembles of weather forecasts at scale at a small fraction of the cost of traditional physics-based forecasting models. This technology opens up novel opportunities for weather and climate science, and it represents one of the first applications to weather and climate forecasting of probabilistic diffusion models, a generative AI technology behind recent advances in media generation. The need for probabilistic forecasts: the butterfly effect American Association for the Advancement of Science meeting in Washington, D.C., MIT meteorology professor Ed Lorenz gave a talk entitled, “Does the Flap of a Butterfly's Wings in Brazil Set Off a Tornado in Texas?” which contributed to the term “butterfly effect”. He was building on his earlier, landmark 1963 paper where he examined the feasibility of “very-long-range weather prediction” and described how errors in initial conditions grow exponentially when integrated in time with numerical weather prediction models. This exponential error growth, known as chaos, results in a deterministic predictability limit that restricts the use of individual forecasts in decision making, because they do not quantify the inherent uncertainty of weather conditions. This is particularly problematic when forecasting extreme weather events, such as hurricanes, heatwaves, or floods. Recognizing the limitations of deterministic forecasts, weather agencies around the world issue probabilistic forecasts. Such forecasts are based on ensembles of deterministic forecasts, each of which is generated by including synthetic noise in the initial conditions and stochasticity in the physical processes. Leveraging the fast error growth rate in weather models, the forecasts in an ensemble are purposefully different: the initial uncertainties are tuned to generate runs that are as different as possible and the stochastic processes in the weather model introduce additional differences during the model run. The error growth is mitigated by averaging all the forecasts in the ensemble and the variability in the ensemble of forecasts quantifies the uncertainty of the weather conditions. While effective, generating these probabilistic forecasts is computationally costly. They require running highly complex numerical weather models on massive supercomputers multiple times. Consequently, many operational weather forecasts can only afford to generate ~10–50 ensemble members for each forecast cycle. This is a problem for users concerned with the likelihood of rare but high-impact weather events, which typically require much larger ensembles to assess beyond a few days. For instance, one would need a 10,000-member ensemble to forecast the likelihood of events with 1% probability of occurrence with a relative error less than 10%. Quantifying the probability of such extreme events could be useful, for example, for emergency management preparation or for energy traders. SEEDS: AI-enabled advances paper, we present the Scalable Ensemble Envelope Diffusion Sampler (SEEDS), a generative AI technology for weather forecast ensemble generation. SEEDS is based on denoising diffusion probabilistic models, a state-of-the-art generative AI method pioneered in part by Google Research. SEEDS can generate a large ensemble conditioned on as few as one or two forecasts from an operational numerical weather prediction system. The generated ensembles not only yield plausible real-weather–like forecasts but also match or exceed physics-based ensembles in skill metrics such as the rank histogram, the root-mean-squared error (RMSE), and the continuous ranked probability score (CRPS). In particular, the generated ensembles assign more accurate likelihoods to the tail of the forecast distribution, such as ±2σ and ±3σ weather events. Most importantly, the computational cost of the model is negligible when compared to the hours of computational time needed by supercomputers to make a forecast. It has a throughput of 256 ensemble members (at 2° resolution) per 3 minutes on Google Cloud TPUv3-32 instances and can easily scale to higher throughput by deploying more accelerators. SEEDS generates an order-of-magnitude more samples to in-fill distributions of weather patterns. Generating plausible weather forecasts Global Ensemble Forecast System, GEFS) for a particular date during the 2022 European heat waves. We also compare the results to the forecasts from a Gaussian model that predicts the univariate mean and standard deviation of each atmospheric field at each location, a common and computationally efficient but less sophisticated data-driven approach. This Gaussian model is meant to characterize the output of pointwise post-processing, which ignores correlations and treats each grid point as an independent random variable. In contrast, a real weather map would have detailed correlational structures. Because SEEDS directly models the joint distribution of the atmospheric state, it realistically captures both the spatial covariance and the correlation between mid-tropospheric geopotential and mean sea level pressure, both of which are closely related and are commonly used by weather forecasters for evaluation and verification of forecasts. Gradients in the mean sea level pressure are what drive winds at the surface, while gradients in mid-tropospheric geopotential create upper-level winds that move large-scale weather patterns. The generated samples from SEEDS shown in the figure below (frames Ca–Ch) display a geopotential trough west of Portugal with spatial structure similar to that found in the operational U.S. forecasts or the reanalysis based on observations. Although the Gaussian model predicts the marginal univariate distributions adequately, it fails to capture cross-field or spatial correlations. This hinders the assessment of the effects that these anomalies may have on hot air intrusions from North Africa, which can exacerbate heat waves over Europe. Stamp maps over Europe on 2022/07/14 at 0:00 UTC. The contours are for the mean sea level pressure (dashed lines mark isobars below 1010 hPa) while the heatmap depicts the geopotential height at the 500 hPa pressure level. (A) The ERA5 reanalysis, a proxy for real observations. (Ba-Bb) 2 members from the 7-day U.S. operational forecasts used as seeds to our model. (Ca-Ch) 8 samples drawn from SEEDS. (Da-Dh) 8 non-seeding members from the 7-day U.S. operational ensemble forecast. (Ea-Ed) 4 samples from a pointwise Gaussian model parameterized by the mean and variance of the entire U.S. operational ensemble. Covering extreme events more accurately SEEDS provides better statistical coverage of the 2022/07/14 European extreme heat event, denoted by the brown star . Each plot shows the values of the total column-integrated water vapor (TCVW) vs. temperature over a grid point near Lisbon, Portugal from 16,384 samples generated by our models, shown as green dots, conditioned on 2 seeds (blue squares) taken from the 7-day U.S. operational ensemble forecasts (denoted by the sparser brown triangles). The valid forecast time is 1:00 local time. The solid contour levels correspond to iso-proportions of the kernel density of SEEDS, with the outermost one encircling 95% of the mass and 11.875% between each level. Conclusion and future outlook Acknowledgements All SEEDS authors, Lizao Li, Rob Carver, Ignacio Lopez-Gomez, Fei Sha and John Anderson, co-authored this blog post, with Carla Bromberg as Program Lead. We also thank Tom Small who designed the animation. Our colleagues at Google Research have provided invaluable advice to the SEEDS work. Among them, we thank Leonardo Zepeda-Núñez, Zhong Yi Wan, Stephan Rasp, Stephan Hoyer, and Tapio Schneider for their inputs and useful discussion. We thank Tyler Russell for additional technical program management, as well as Alex Merose for data coordination and support. We also thank Cenk Gazen, Shreya Agrawal, and Jason Hickey for discussions in the early stage of the SEEDS work.

    AI / MLopen article
  • Google AI Blog28/03/2024, 20:53

    AutoBNN: Probabilistic time series forecasting with compositional bayesian neural networks

    Posted by Urs Köster, Software Engineer, Google Research Time series problems are ubiquitous, from forecasting weather and traffic patterns to understanding economic trends. Bayesian approaches start with an assumption about the data's patterns (prior probability), collecting evidence (e.g., new time series data), and continuously updating that assumption to form a posterior probability distribution. Traditional Bayesian approaches like Gaussian processes (GPs) and Structural Time Series are extensively used for modeling time series data, e.g., the commonly used Mauna Loa CO2 dataset. However, they often rely on domain experts to painstakingly select appropriate model components and may be computationally expensive. Alternatives such as neural networks lack interpretability, making it difficult to understand how they generate forecasts, and don't produce reliable confidence intervals. To that end, we introduce AutoBNN, a new open-source package written in JAX. AutoBNN automates the discovery of interpretable time series forecasting models, provides high-quality uncertainty estimates, and scales effectively for use on large datasets. We describe how AutoBNN combines the interpretability of traditional probabilistic approaches with the scalability and flexibility of neural networks. AutoBNN line of research that over the past decade has yielded improved predictive accuracy by modeling time series using GPs with learned kernel structures. The kernel function of a GP encodes assumptions about the function being modeled, such as the presence of trends, periodicity or noise. With learned GP kernels, the kernel function is defined compositionally: it is either a base kernel (such as Linear, Quadratic, Periodic, Matérn or ExponentiatedQuadratic) or a composite that combines two or more kernel functions using operators such as Addition, Multiplication, or ChangePoint. This compositional kernel structure serves two related purposes. First, it is simple enough that a user who is an expert about their data, but not necessarily about GPs, can construct a reasonable prior for their time series. Second, techniques like Sequential Monte Carlo can be used for discrete searches over small structures and can output interpretable results. Bayesian neural networks (BNNs) while retaining the compositional kernel structure. A BNN is a neural network with a probability distribution over weights rather than a fixed set of weights. This induces a distribution over outputs, capturing uncertainty in the predictions. BNNs bring the following advantages over GPs: First, training large GPs is computationally expensive, and traditional training algorithms scale as the cube of the number of data points in the time series. In contrast, for a fixed width, training a BNN will often be approximately linear in the number of data points. Second, BNNs lend themselves better to GPU and TPU hardware acceleration than GP training operations. Third, compositional BNNs can be easily combined with traditional deep BNNs, which have the ability to do feature discovery. One could imagine "hybrid" architectures, in which users specify a top-level structure of Add(Linear, Periodic, Deep), and the deep BNN is left to learn the contributions from potentially high-dimensional covariate information. How might one translate a GP with compositional kernels into a BNN then? A single layer neural network will typically converge to a GP as the number of neurons (or "width") goes to infinity. More recently, researchers have discovered a correspondence in the other direction — many popular GP kernels (such as Matern, ExponentiatedQuadratic, Polynomial or Periodic) can be obtained as infinite-width BNNs with appropriately chosen activation functions and weight distributions. Furthermore, these BNNs remain close to the corresponding GP even when the width is very much less than infinite. For example, the figures below show the difference in the covariance between pairs of observations, and regression results of the true GPs and their corresponding width-10 neural network versions. Comparison of Gram matrices between true GP kernels (top row) and their width 10 neural network approximations (bottom row). Comparison of regression results between true GP kernels (top row) and their width 10 neural network approximations (bottom row). BNN analogues of the Addition and Multiplication operators over GPs, and input warping to produce periodic kernels. BNN addition is straightforwardly given by adding the outputs of the component BNNs. BNN multiplication is achieved by multiplying the activations of the hidden layers of the BNNs and then applying a shared dense layer. We are therefore limited to only multiplying BNNs with the same hidden width. Using AutoBNN package is available within Tensorflow Probability. It is implemented in JAX and uses the flax.linen neural network library. It implements all of the base kernels and operators discussed so far (Linear, Quadratic, Matern, ExponentiatedQuadratic, Periodic, Addition, Multiplication) plus one new kernel and three new operators: a OneLayer kernel, a single hidden layer ReLU BNN, a ChangePoint operator that allows smoothly switching between two kernels, a LearnableChangePoint operator which is the same as ChangePoint except position and slope are given prior distributions and can be learnt from the data, and a WeightedSum operator. WeightedSum combines two or more BNNs with learnable mixing weights, where the learnable weights follow a Dirichlet prior. By default, a flat Dirichlet distribution with concentration 1.0 is used. WeightedSums allow a "soft" version of structure discovery, i.e., training a linear combination of many possible models at once. In contrast to structure discovery with discrete structures, such as in AutoGP, this allows us to use standard gradient methods to learn structures, rather than using expensive discrete optimization. Instead of evaluating potential combinatorial structures in series, WeightedSum allows us to evaluate them in parallel. To easily enable exploration, AutoBNN defines a number of model structures that contain either top-level or internal WeightedSums. The names of these models can be used as the first parameter in any of the estimator constructors, and include things like sum_of_stumps (the WeightedSum over all the base kernels) and sum_of_shallow (which adds all possible combinations of base kernels with all operators). Illustration of the sum_of_stumps model. The bars in the top row show the amount by which each base kernel contributes, and the bottom row shows the function represented by the base kernel. The resulting weighted sum is shown on the right. M3 dataset. The six base structures were ExponentiatedQuadratic (which is the same as the Radial Basis Function kernel, or RBF for short), Matern, Linear, Quadratic, OneLayer and Periodic kernels. The figure shows the MAP estimates of their weights over an ensemble of 32 particles. All of the high likelihood particles gave a large weight to the Periodic component, low weights to Linear, Quadratic and OneLayer, and a large weight to either RBF or Matern. Parallel coordinates plot of the MAP estimates of the base kernel weights over 32 particles. The sum_of_stumps model was trained on the N374 series from the M3 dataset (insert in blue). Darker lines correspond to particles with higher likelihoods. WeightedSums as the inputs to other operators, it is possible to express rich combinatorial structures, while keeping models compact and the number of learnable weights small. As an example, we include the sum_of_products model (illustrated in the figure below) which first creates a pairwise product of two WeightedSums, and then a sum of the two products. By setting some of the weights to zero, we can create many different discrete structures. The total number of possible structures in this model is 216, since there are 16 base kernels that can be turned on or off. All these structures are explored implicitly by training just this one model. Illustration of the "sum_of_products" model. Each of the four WeightedSums have the same structure as the "sum_of_stumps" model. Periodic and either the Matern or ExponentiatedQuadratic) lead to overfitting on many datasets. To prevent this, we have defined model classes like sum_of_safe_shallow that exclude such products when performing structure discovery with WeightedSums. For training, AutoBNN provides AutoBnnMapEstimator and AutoBnnMCMCEstimator to perform MAP and MCMC inference, respectively. Either estimator can be combined with any of the six likelihood functions, including four based on normal distributions with different noise characteristics for continuous data and two based on the negative binomial distribution for count data. Result from running AutoBNN on the Mauna Loa CO2 dataset in our example colab. The model captures the trend and seasonal component in the data. Extrapolating into the future, the mean prediction slightly underestimates the actual trend, while the 95% confidence interval gradually increases. scikit-learn–inspired estimator interface: import autobnn as ab model = ab.operators.Add( bnns=(ab.kernels.PeriodicBNN(width=50), ab.kernels.LinearBNN(width=50), ab.kernels.MaternBNN(width=50))) estimator = ab.estimators.AutoBnnMapEstimator( model, 'normal_likelihood_logistic_noise', jax.random.PRNGKey(42), periods=[12]) estimator.fit(my_training_data_xs, my_training_data_ys) low, mid, high = estimator.predict_quantiles(my_training_data_xs) Conclusion AutoBNN provides a powerful and flexible framework for building sophisticated time series prediction models. By combining the strengths of BNNs and GPs with compositional kernels, AutoBNN opens a world of possibilities for understanding and forecasting complex data. We invite the community to try the colab, and leverage this library to innovate and solve real-world challenges. Acknowledgements AutoBNN was written by Colin Carroll, Thomas Colthurst, Urs Köster and Srinivas Vasudevan. We would like to thank Kevin Murphy, Brian Patton and Feras Saad for their advice and feedback.

    AI / MLopen article
  • Google AI Blog20/03/2024, 20:54

    Computer-aided diagnosis for lung cancer screening

    Posted by Atilla Kiraly, Software Engineer, and Rory Pilgrim, Product Manager, Google Research Lung cancer is the leading cause of cancer-related deaths globally with 1.8 million deaths reported in 2020. Late diagnosis dramatically reduces the chances of survival. Lung cancer screening via computed tomography (CT), which provides a detailed 3D image of the lungs, has been shown to reduce mortality in high-risk populations by at least 20% by detecting potential signs of cancers earlier. In the US, screening involves annual scans, with some countries or cases recommending more or less frequent scans. The United States Preventive Services Task Force recently expanded lung cancer screening recommendations by roughly 80%, which is expected to increase screening access for women and racial and ethnic minority groups. However, false positives (i.e., incorrectly reporting a potential cancer in a cancer-free patient) can cause anxiety and lead to unnecessary procedures for patients while increasing costs for the healthcare system. Moreover, efficiency in screening a large number of individuals can be challenging depending on healthcare infrastructure and radiologist availability. At Google we have previously developed machine learning (ML) models for lung cancer detection, and have evaluated their ability to automatically detect and classify regions that show signs of potential cancer. Performance has been shown to be comparable to that of specialists in detecting possible cancer. While they have achieved high performance, effectively communicating findings in realistic environments is necessary to realize their full potential. To that end, in “Assistive AI in Lung Cancer Screening: A Retrospective Multinational Study in the US and Japan”, published in Radiology AI, we investigate how ML models can effectively communicate findings to radiologists. We also introduce a generalizable user-centric interface to help radiologists leverage such models for lung cancer screening. The system takes CT imaging as input and outputs a cancer suspicion rating using four categories (no suspicion, probably benign, suspicious, highly suspicious) along with the corresponding regions of interest. We evaluate the system’s utility in improving clinician performance through randomized reader studies in both the US and Japan, using the local cancer scoring systems (Lung-RADSs V1.1 and Sendai Score) and image viewers that mimic realistic settings. We found that reader specificity increases with model assistance in both reader studies. To accelerate progress in conducting similar studies with ML models, we have open-sourced code to process CT images and generate images compatible with the picture archiving and communication system (PACS) used by radiologists. Developing an interface to communicate model results alpha-numeric score to indicate the lung cancer risk and follow-up recommendations. When assessing patients, radiologists load the CT in their workstation to read the case, find lung nodules or lesions, and apply set guidelines to determine follow-up decisions. Our first step was to improve the previously developed ML models through additional training data and architectural improvements, including self-attention. Then, instead of targeting specific guidelines, we experimented with a complementary way of communicating AI results independent of guidelines or their particular versions. Specifically, the system output offers a suspicion rating and localization (regions of interest) for the user to consider in conjunction with their own specific guidelines. The interface produces output images directly associated with the CT study, requiring no changes to the user’s workstation. The radiologist only needs to review a small set of additional images. There is no other change to their system or interaction with the system. Example of the assistive lung cancer screening system outputs. Results for the radiologist’s evaluation are visualized on the location of the CT volume where the suspicious lesion is found. The overall suspicion is displayed at the top of the CT images. Circles highlight the suspicious lesions while squares show a rendering of the same lesion from a different perspective, called a sagittal view. prior work. The models coordinate with each other to first segment the lungs, obtain an overall assessment, locate three suspicious regions, then use the information to assign a suspicion rating to each region. The system was deployed on Google Cloud using a Google Kubernetes Engine (GKE) that pulled the images, ran the ML models, and provided results. This allows scalability and directly connects to servers where the images are stored in DICOM stores. Outline of the Google Cloud deployment of the assistive lung cancer screening system and the directional calling flow for the individual components that serve the images and compute results. Images are served to the viewer and to the system using Google Cloud services. The system is run on a Google Kubernetes Engine that pulls the images, processes them, and writes them back into the DICOM store. Reader studies area under the ROC curve (AUC) values. These were compared with and without assistance. A multi-case multi-reader study involves each case being reviewed by each reader twice, once with ML system assistance and once without. In this visualization one reader first reviews Set A without assistance (blue) and then with assistance (orange) after a wash-out period. A second reader group follows the opposite path by reading the same set of cases Set A with assistance first. Readers are randomized to these groups to remove the effect of ordering. specificity) by an absolute 5–7% compared to when they didn’t use the assistive system. This potentially means that for every 15–20 patients screened, one may be able to avoid unnecessary follow-up procedures, thus reducing their anxiety and the burden on the health care system. This can, in turn, help improve the sustainability of lung cancer screening programs, particularly as more people become eligible for screening. Reader specificity increases with ML model assistance in both the US-based and Japan-based reader studies. Specificity values were derived from reader scores from actionable findings (something suspicious was found) versus no actionable findings, compared against the true cancer outcome of the individual. Under model assistance, readers flagged fewer cancer-negative individuals for follow-up visits. Sensitivity for cancer positive individuals remained the same. Translating this into real-world impact through partnership DeepHealth, a leading AI-powered health informatics provider; and Apollo Radiology International a leading provider of Radiology services in India to explore paths for incorporating this system into future products. In addition, we are looking to help other researchers studying how best to integrate ML model results into clinical workflows by open sourcing code used for the reader study and incorporating the insights described in this blog. We hope that this will help accelerate medical imaging researchers looking to conduct reader studies for their AI models, and catalyze translational research in the field. Acknowledgements Key contributors to this project include Corbin Cunningham, Zaid Nabulsi, Ryan Najafi, Jie Yang, Charles Lau, Joseph R. Ledsam, Wenxing Ye, Diego Ardila, Scott M. McKinney, Rory Pilgrim, Hiroaki Saito, Yasuteru Shimamura, Mozziyar Etemadi, Yun Liu, David Melnick, Sunny Jansen, Nadia Harhen, David P. Nadich, Mikhail Fomitchev, Ziyad Helali, Shabir Adeel, Greg S. Corrado, Lily Peng, Daniel Tse, Shravya Shetty, Shruthi Prabhakara, Neeral Beladia, and Krish Eswaran. Thanks to Arnav Agharwal and Andrew Sellergren for their open sourcing support and Vivek Natarajan and Michael D. Howell for their feedback. Sincere appreciation also goes to the radiologists who enabled this work with their image interpretation and annotation efforts throughout the study, and Jonny Wong and Carli Sampson for coordinating the reader studies.

    AI / MLopen article
  • Google AI Blog20/03/2024, 16:06

    Using AI to expand global access to reliable flood forecasts

    Posted by Yossi Matias, VP Engineering & Research, and Grey Nearing, Research Scientist, Google Research Floods are the most common natural disaster, and are responsible for roughly $50 billion in annual financial damages worldwide. The rate of flood-related disasters has more than doubled since the year 2000 partly due to climate change. Nearly 1.5 billion people, making up 19% of the world’s population, are exposed to substantial risks from severe flood events. Upgrading early warning systems to make accurate and timely information accessible to these populations can save thousands of lives per year. Driven by the potential impact of reliable flood forecasting on people’s lives globally, we started our flood forecasting effort in 2017. Through this multi-year journey, we advanced research over the years hand-in-hand with building a real-time operational flood forecasting system that provides alerts on Google Search, Maps, Android notifications and through the Flood Hub. However, in order to scale globally, especially in places where accurate local data is not available, more research advances were required. In “Global prediction of extreme floods in ungauged watersheds”, published in Nature, we demonstrate how machine learning (ML) technologies can significantly improve global-scale flood forecasting relative to the current state-of-the-art for countries where flood-related data is scarce. With these AI-based technologies we extended the reliability of currently-available global nowcasts, on average, from zero to five days, and improved forecasts across regions in Africa and Asia to be similar to what are currently available in Europe. The evaluation of the models was conducted in collaboration with the European Center for Medium Range Weather Forecasting (ECMWF). These technologies also enable Flood Hub to provide real-time river forecasts up to seven days in advance, covering river reaches across over 80 countries. This information can be used by people, communities, governments and international organizations to take anticipatory action to help protect vulnerable populations. Flood forecasting at Google launched a pilot early warning system in the Ganges-Brahmaputra river basin in India, with the hypothesis that ML could help address the challenging problem of reliable flood forecasting at scale. The pilot was further expanded the following year via the combination of an inundation model, real-time water level measurements, the creation of an elevation map and hydrologic modeling. In collaboration with academics, and, in particular, with the JKU Institute for Machine Learning we explored ML-based hydrologic models, showing that LSTM-based models could produce more accurate simulations than traditional conceptual and physics-based hydrology models. This research led to flood forecasting improvements that enabled the expansion of our forecasting coverage to include all of India and Bangladesh. We also worked with researchers at Yale University to test technological interventions that increase the reach and impact of flood warnings. Our hydrological models predict river floods by processing publicly available weather data like precipitation and physical watershed information. Such models must be calibrated to long data records from streamflow gauging stations in individual rivers. A low percentage of global river watersheds (basins) have streamflow gauges, which are expensive but necessary to supply relevant data, and it’s challenging for hydrological simulation and forecasting to provide predictions in basins that lack this infrastructure. Lower gross domestic product (GDP) is correlated with increased vulnerability to flood risks, and there is an inverse correlation between national GDP and the amount of publicly available data in a country. ML helps to address this problem by allowing a single model to be trained on all available river data and to be applied to ungauged basins where no data are available. In this way, models can be trained globally, and can make predictions for any river location. There is an inverse (log-log) correlation between the amount of publicly available streamflow data in a country and national GDP. Streamflow data from the Global Runoff Data Center. estimate uncertainty in river forecasts and showed how ML river forecast models synthesize information from multiple data sources. They demonstrated that these models can simulate extreme events reliably, even when those events are not part of the training data. In an effort to contribute to open science, in 2023 we open-sourced a community-driven dataset for large-sample hydrology in Nature Scientific Data. The river forecast model LSTMs perform well on the task of river forecasting. A diagram of the LSTM, which is a neural network that operates sequentially in time. An accessible primer can be found here. mixture density networks to produce a probabilistic forecast (i.e., predicted parameters of a probability distribution over streamflow). Specifically, the model predicts the parameters of a mixture of heavy-tailed probability density functions, called asymmetric Laplacian distributions, at each forecast time step. The result is a mixture density function, called a Countable Mixture of Asymmetric Laplacians (CMAL) distribution, which represents a probabilistic prediction of the volumetric flow rate in a particular river at a particular time. LSTM-based river forecast model architecture. Two LSTMs are applied in sequence, one ingesting historical weather data and one ingesting forecasted weather data. The model outputs are the parameters of a probability distribution over streamflow at each forecasted timestep. Input and training data Static watershed attributes representing geographical and geophysical variables: From the HydroATLAS project, including data like long-term climate indexes (precipitation, temperature, snow fractions), land cover, and anthropogenic attributes (e.g., a nighttime lights index as a proxy for human development). Historical meteorological time-series data: Used to spin up the model for one year prior to the issue time of a forecast. The data comes from NASA IMERG, NOAA CPC Global Unified Gauge-Based Analysis of Daily Precipitation, and the ECMWF ERA5-land reanalysis. Variables include daily total precipitation, air temperature, solar and thermal radiation, snowfall, and surface pressure. Forecasted meteorological time series over a seven-day forecast horizon: Used as input for the forecast LSTM. These data are the same meteorological variables listed above, and come from the ECMWF HRES atmospheric model. Training data are daily streamflow values from the Global Runoff Data Center over the time period 1980 - 2023. A single streamflow forecast model is trained using data from 5,680 diverse watershed streamflow gauges (shown below) to improve accuracy. Location of 5,680 streamflow gauges that supply training data for the river forecast model from the Global Runoff Data Center. Improving on the current state-of-the-art GloFAS version 4, the current state-of-the-art global flood forecasting system. These experiments showed that ML can provide accurate warnings earlier and over larger and more impactful events. The figure below shows the distribution of F1 scores when predicting different severity events at river locations around the world, with plus or minus 1 day accuracy. F1 scores are an average of precision and recall and event severity is measured by return period. For example, a 2-year return period event is a volume of streamflow that is expected to be exceeded on average once every two years. Our model achieves reliability scores at up to 4-day or 5-day lead times that are similar to or better, on average, than the reliability of GloFAS nowcasts (0-day lead time). Distributions of F1 scores over 2-year return period events in 2,092 watersheds globally during the time period 2014-2023 from GloFAS (blue) and our model (orange) at different lead times. On average, our model is statistically as accurate as GloFAS nowcasts (0–day lead time) up to 5 days in advance over 2-year (shown) and 1-year, 5-year, and 10-year events (not shown). paper for more information. Looking into the future Adaptation and Resilience efforts and reflects Google's commitment to address climate change while helping global communities become more resilient. We believe that AI and ML will continue to play a critical role in helping advance science and research towards climate action. We actively collaborate with several international aid organizations (e.g., the Centre for Humanitarian Data and the Red Cross) to provide actionable flood forecasts. Additionally, in an ongoing collaboration with the World Meteorological Organization (WMO) to support early warning systems for climate hazards, we are conducting a study to help understand how AI can help address real-world challenges faced by national flood forecasting agencies. While the work presented here demonstrates a significant step forward in flood forecasting, future work is needed to further expand flood forecasting coverage to more locations globally and other types of flood-related events and disasters, including flash floods and urban floods. We are looking forward to continuing collaborations with our partners in the academic and expert communities, local governments and the industry to reach these goals.

    AI / MLopen article
  • Google AI Blog19/03/2024, 20:15

    ScreenAI: A visual language model for UI and visually-situated language understanding

    Posted by Srinivas Sunkara and Gilles Baechler, Software Engineers, Google Research Screen user interfaces (UIs) and infographics, such as charts, diagrams and tables, play important roles in human communication and human-machine interaction as they facilitate rich and interactive user experiences. UIs and infographics share similar design principles and visual language (e.g., icons and layouts), that offer an opportunity to build a single model that can understand, reason, and interact with these interfaces. However, because of their complexity and varied presentation formats, infographics and UIs present a unique modeling challenge. To that end, we introduce “ScreenAI: A Vision-Language Model for UI and Infographics Understanding”. ScreenAI improves upon the PaLI architecture with the flexible patching strategy from pix2struct. We train ScreenAI on a unique mixture of datasets and tasks, including a novel Screen Annotation task that requires the model to identify UI element information (i.e., type, location and description) on a screen. These text annotations provide large language models (LLMs) with screen descriptions, enabling them to automatically generate question-answering (QA), UI navigation, and summarization training datasets at scale. At only 5B parameters, ScreenAI achieves state-of-the-art results on UI- and infographic-based tasks (WebSRC and MoTIF), and best-in-class performance on Chart QA, DocVQA, and InfographicVQA compared to models of similar size. We are also releasing three new datasets: Screen Annotation to evaluate the layout understanding capability of the model, as well as ScreenQA Short and Complex ScreenQA for a more comprehensive evaluation of its QA capability. ScreenAI PaLI, composed of a multimodal encoder block and an autoregressive decoder. The PaLI encoder uses a vision transformer (ViT) that creates image embeddings and a multimodal encoder that takes the concatenation of the image and text embeddings as input. This flexible architecture allows ScreenAI to solve vision tasks that can be recast as text+image-to-text problems. On top of the PaLI architecture, we employ a flexible patching strategy introduced in pix2struct. Instead of using a fixed-grid pattern, the grid dimensions are selected such that they preserve the native aspect ratio of the input image. This enables ScreenAI to work well across images of various aspect ratios. The ScreenAI model is trained in two stages: a pre-training stage followed by a fine-tuning stage. First, self-supervised learning is applied to automatically generate data labels, which are then used to train ViT and the language model. ViT is frozen during the fine-tuning stage, where most data used is manually labeled by human raters. ScreenAI model architecture. Data generation publicly accessible web pages and following the programmatic exploration approach used for the RICO dataset for mobile apps. We then apply a layout annotator, based on the DETR model, that identifies and labels a wide range of UI elements (e.g., image, pictogram, button, text) and their spatial relationships. Pictograms undergo further analysis using an icon classifier capable of distinguishing 77 different icon types. This detailed classification is essential for interpreting the subtle information conveyed through icons. For icons that are not covered by the classifier, and for infographics and images, we use the PaLI image captioning model to generate descriptive captions that provide contextual information. We also apply an optical character recognition (OCR) engine to extract and annotate textual content on screen. We combine the OCR text with the previous annotations to create a detailed description of each screen. A mobile app screenshot with generated annotations that include UI elements and their descriptions, e.g., TEXT elements also contain the text content from OCR, IMAGE elements contain image captions, LIST_ITEMs contain all their child elements. LLM-based data generation PaLM 2 to generate input-output pairs in a two-step process. First, screen annotations are generated using the technique outlined above, then we craft a prompt around this schema for the LLM to create synthetic data. This process requires prompt engineering and iterative refinement to find an effective prompt. We assess the generated data's quality through human validation against a quality threshold. You only speak JSON. Do not write text that isn’t JSON. You are given the following mobile screenshot, described in words. Can you generate 5 questions regarding the content of the screenshot as well as the corresponding short answers to them? The answer should be as short as possible, containing only the necessary information. Your answer should be structured as follows: questions: [ {{question: the question, answer: the answer }}, ... ] {THE SCREEN SCHEMA} A sample prompt for QA data generation. Question answering: The model is asked to answer questions regarding the content of the screenshots, e.g., “When does the restaurant open?” Screen navigation: The model is asked to convert a natural language utterance into an executable action on a screen, e.g., “Click the search button.” Screen summarization: The model is asked to summarize the screen content in one or two sentences. Block diagram of our workflow for generating data for QA, summarization and navigation tasks using existing ScreenAI models and LLMs. Each task uses a custom prompt to emphasize desired aspects, like questions related to counting, involving reasoning, etc. LLM-generated data. Examples for screen QA, navigation and summarization. For navigation, the action bounding box is displayed in red on the screenshot. Experiments and results ChartQA, DocVQA, Multi page DocVQA, InfographicVQA, OCR VQA, Web SRC and ScreenQA. For navigation, datasets used include Referring Expressions, MoTIF, Mug, and Android in the Wild. Finally, we use Screen2Words for screen summarization and Widget Captioning for describing specific UI elements. Along with the fine-tuning datasets, we evaluate the fine-tuned ScreenAI model using three novel benchmarks: Screen Annotation: Enables the evaluation model layout annotations and spatial understanding capabilities. ScreenQA Short: A variation of ScreenQA, where its ground truth answers have been shortened to contain only the relevant information that better aligns with other QA tasks. Complex ScreenQA: Complements ScreenQA Short with more difficult questions (counting, arithmetic, comparison, and non-answerable questions) and contains screens with various aspect ratios. The fine-tuned ScreenAI model achieves state-of-the-art results on various UI and infographic-based tasks (WebSRC and MoTIF) and best-in-class performance on Chart QA, DocVQA, and InfographicVQA compared to models of similar size. ScreenAI achieves competitive performance on Screen2Words and OCR-VQA. Additionally, we report results on the new benchmark datasets introduced to serve as a baseline for further research. Comparing model performance of ScreenAI with state-of-the-art (SOTA) models of similar size. Model performance increases with size, and the performance has not saturated even at the largest size of 5B params. Conclusion Acknowledgements This project is the result of joint work with Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Carbune, Jason Lin, Jindong Chen and Abhanshu Sharma. We thank Fangyu Liu, Xi Chen, Efi Kokiopoulou, Jesse Berent, Gabriel Barcik, Lukas Zilka, Oriana Riva, Gang Li,Yang Li, Radu Soricut, and Tania Bedrax-Weiss for their insightful feedback and discussions, along with Rahul Aralikatte, Hao Cheng and Daniel Kim for their support in data preparation. We also thank Jay Yagnik, Blaise Aguera y Arcas, Ewa Dominowska, David Petrou, and Matt Sharifi for their leadership, vision and support. We are very grateful toTom Small for helping us create the animation in this post.

    AI / MLopen article
  • Google AI Blog19/03/2024, 15:00

    SCIN: A new resource for representative dermatology images

    Posted by Pooja Rao, Research Scientist, Google Research Health datasets play a crucial role in research and medical education, but it can be challenging to create a dataset that represents the real world. For example, dermatology conditions are diverse in their appearance and severity and manifest differently across skin tones. Yet, existing dermatology image datasets often lack representation of everyday conditions (like rashes, allergies and infections) and skew towards lighter skin tones. Furthermore, race and ethnicity information is frequently missing, hindering our ability to assess disparities or create solutions. To address these limitations, we are releasing the Skin Condition Image Network (SCIN) dataset in collaboration with physicians at Stanford Medicine. We designed SCIN to reflect the broad range of concerns that people search for online, supplementing the types of conditions typically found in clinical datasets. It contains images across various skin tones and body parts, helping to ensure that future AI tools work effectively for all. We've made the SCIN dataset freely available as an open-access resource for researchers, educators, and developers, and have taken careful steps to protect contributor privacy. Example set of images and metadata from the SCIN dataset. Dataset composition tanning propensity (self-reported Fitzpatrick Skin Type, i.e., sFST), and to describe the texture, duration and symptoms related to their concern. One to three dermatologists labeled each contribution with up to five dermatology conditions, along with a confidence score for each label. The SCIN dataset contains these individual labels, as well as an aggregated and weighted differential diagnosis derived from them that could be useful for model testing or training. These labels were assigned retrospectively and are not equivalent to a clinical diagnosis, but they allow us to compare the distribution of dermatology conditions in the SCIN dataset with existing datasets. The SCIN dataset contains largely allergic, inflammatory and infectious conditions while datasets from clinical sources focus on benign and malignant neoplasms. Monk Skin Tone (eMST) for the images. This allowed comparison of the skin condition and skin type distributions to those in existing dermatology datasets. Although we did not selectively target any skin types or skin tones, the SCIN dataset has a balanced Fitzpatrick skin type distribution (with more of Types 3, 4, 5, and 6) compared to similar datasets from clinical sources. Self-reported and dermatologist-estimated Fitzpatrick Skin Type distribution in the SCIN dataset compared with existing un-enriched dermatology datasets (Fitzpatrick17k, PH², SKINL2, and PAD-UFES-20). Fitzpatrick Skin Type scale was originally developed as a photo-typing scale to measure the response of skin types to UV radiation, and it is widely used in dermatology research. The Monk Skin Tone scale is a newer 10-shade scale that measures skin tone rather than skin phototype, capturing more nuanced differences between the darker skin tones. While neither scale was intended for retrospective estimation using images, the inclusion of these labels is intended to enable future research into skin type and tone representation in dermatology. For example, the SCIN dataset provides an initial benchmark for the distribution of these skin types and tones in the US population. The SCIN dataset has a high representation of women and younger individuals, likely reflecting a combination of factors. These could include differences in skin condition incidence, propensity to seek health information online, and variations in willingness to contribute to research across demographics. Crowdsourcing method research paper co-authored with investigators at Stanford Medicine. This approach empowers individuals to play an active role in healthcare research. It allows us to reach people at earlier stages of their health concerns, potentially before they seek formal care. Crucially, this method uses advertisements on web search result pages — the starting point for many people’s health journey — to connect with participants. Our results demonstrate that crowdsourcing can yield a high-quality dataset with a low spam rate. Over 97.5% of contributions were genuine images of skin conditions. After performing further filtering steps to exclude images that were out of scope for the SCIN dataset and to remove duplicates, we were able to release nearly 90% of the contributions received over the 8-month study period. Most images were sharp and well-exposed. Approximately half of the contributions include self-reported demographics, and 80% contain self-reported information relating to the skin condition, such as texture, duration, or other symptoms. We found that dermatologists’ ability to retrospectively assign a differential diagnosis depended more on the availability of self-reported information than on image quality. Dermatologist confidence in their labels (scale from 1-5) depended on the availability of self-reported demographic and symptom information. Data Use License prohibits attempts to re-identify contributors. We hope the SCIN dataset will be a helpful resource for those working to advance inclusive dermatology research, education, and AI tool development. By demonstrating an alternative to traditional dataset creation methods, SCIN paves the way for more representative datasets in areas where self-reported data or retrospective labeling is feasible. Acknowledgements We are grateful to all our co-authors Abbi Ward, Jimmy Li, Julie Wang, Sriram Lakshminarasimhan, Ashley Carrick, Bilson Campana, Jay Hartford, Pradeep Kumar S, Tiya Tiyasirisokchai, Sunny Virmani, Renee Wong, Yossi Matias, Greg S. Corrado, Dale R. Webster, Dawn Siegel (Stanford Medicine), Steven Lin (Stanford Medicine), Justin Ko (Stanford Medicine), Alan Karthikesalingam and Christopher Semturs. We also thank Yetunde Ibitoye, Sami Lachgar, Lisa Lehmann, Javier Perez, Margaret Ann Smith (Stanford Medicine), Rachelle Sico, Amit Talreja, Annisah Um’rani and Wayne Westerlind for their essential contributions to this work. Finally, we are grateful to Heather Cole-Lewis, Naama Hammel, Ivor Horn, Michael Howell, Yun Liu, and Eric Teasley for their insightful comments on the study design and manuscript.

    AI / MLopen article
  • Google AI Blog18/03/2024, 18:41

    MELON: Reconstructing 3D objects from images with unknown poses

    Posted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object looks like in whole, even if only looking at a few 2D pictures of it. Yet the capacity for a computer to reconstruct the shape of an object in 3D given only a few images has remained a difficult algorithmic problem for years. This fundamental computer vision task has applications ranging from the creation of e-commerce 3D models to autonomous vehicle navigation. A key part of the problem is how to determine the exact positions from which images were taken, known as pose inference. If camera poses are known, a range of successful techniques — such as neural radiance fields (NeRF) or 3D Gaussian Splatting — can reconstruct an object in 3D. But if these poses are not available, then we face a difficult “chicken and egg” problem where we could determine the poses if we knew the 3D object, but we can’t reconstruct the 3D object until we know the camera poses. The problem is made harder by pseudo-symmetries — i.e., many objects look similar when viewed from different angles. For example, square objects like a chair tend to look similar every 90° rotation. Pseudo-symmetries of an object can be revealed by rendering it on a turntable from various angles and plotting its photometric self-similarity map. Self-Similarity map of a toy truck model. Left: The model is rendered on a turntable from various azimuthal angles, θ. Right: The average L2 RGB similarity of a rendering from θ with that of θ*. The pseudo-similarities are indicated by the dashed red lines. ill-posed, with naïve approaches often converging to local minima. In practice, such an approach might mistake the back view as the front view of an object, because they share a similar silhouette. Previous techniques (such as BARF or SAMURAI) side-step this problem by relying on an initial pose estimate that starts close to the global minima. But how can we approach this if those aren’t available? Methods, such as GNeRF and VMRF leverage generative adversarial networks (GANs) to overcome the problem. These techniques have the ability to artificially “amplify” a limited number of training views, aiding reconstruction. GAN techniques, however, often have complex, sometimes unstable, training processes, making robust and reliable convergence difficult to achieve in practice. A range of other successful methods, such as SparsePose or RUST, can infer poses from a limited number views, but require pre-training on a large dataset of posed images, which aren’t always available, and can suffer from “domain-gap” issues when inferring poses for different types of images. In “MELON: NeRF with Unposed Images in SO(3)”, spotlighted at 3DV 2024, we present a technique that can determine object-centric camera poses entirely from scratch while reconstructing the object in 3D. MELON (Modulo Equivalent Latent Optimization of NeRF) is one of the first techniques that can do this without initial pose camera estimates, complex training schemes or pre-training on labeled data. MELON is a relatively simple technique that can easily be integrated into existing NeRF methods. We demonstrate that MELON can reconstruct a NeRF from unposed images with state-of-the-art accuracy while requiring as few as 4–6 images of an object. MELON convolutional neural network (CNN) encoder that regresses camera poses from training images. We pass a downscaled training image to a four layer CNN that infers the camera pose. This CNN is initialized from noise and requires no pre-training. Its capacity is so small that it forces similar looking images to similar poses, providing an implicit regularization greatly aiding convergence. The second technique is a modulo loss that simultaneously considers pseudo symmetries of an object. We render the object from a fixed set of viewpoints for each training image, backpropagating the loss only through the view that best fits the training image. This effectively considers the plausibility of multiple views for each image. In practice, we find N=2 views (viewing an object from the other side) is all that’s required in most cases, but sometimes get better results with N=4 for square objects. These two techniques are integrated into standard NeRF training, except that instead of fixed camera poses, poses are inferred by the CNN and duplicated by the modulo loss. Photometric gradients back-propagate through the best-fitting cameras into the CNN. We observe that cameras generally converge quickly to globally optimal poses (see animation below). After training of the neural field, MELON can synthesize novel views using standard NeRF rendering methods. We simplify the problem by using the NeRF-Synthetic dataset, a popular benchmark for NeRF research and common in the pose-inference literature. This synthetic dataset has cameras at precisely fixed distances and a consistent “up” orientation, requiring us to infer only the polar coordinates of the camera. This is the same as an object at the center of a globe with a camera always pointing at it, moving along the surface. We then only need the latitude and longitude (2 degrees of freedom) to specify the camera pose. MELON uses a dynamically trained lightweight CNN encoder that predicts a pose for each image. Predicted poses are replicated by the modulo loss, which only penalizes the smallest L2 distance from the ground truth color. At evaluation time, the neural field can be used to generate novel views. Results peak signal-to-noise ratio (PSNR) against held out test views. We see that MELON quickly converges to the approximate poses of most cameras within the first 1,000 steps of training, and achieves a competitive PSNR of 27.5 dB after 50k steps. Convergence of MELON on a toy truck model during optimization. Left: Rendering of the NeRF. Right: Polar plot of predicted (blue x), and ground truth (red dot) cameras. Reconstruction quality comparison between ground-truth (GT) and MELON on NeRF-Synthetic scenes after 100k training steps. Noisy images novel view synthesis from extremely noisy, unposed images. We add varying amounts, σ, of white Gaussian noise to the training images. For example, the object in σ=1.0 below is impossible to make out, yet MELON can determine the pose and generate novel views of the object. Novel view synthesis from noisy unposed 128×128 images. Top: Example of noise level present in training views. Bottom: Reconstructed model from noisy training views and mean angular pose error. RawNeRF have demonstrated NeRF’s excellent de-noising capabilities with known camera poses. The fact that MELON works for noisy images of unknown camera poses so robustly was unexpected. Conclusion paper and MELON site to learn more. Acknowledgements We would like to thank our paper co-authors Axel Levy, Matan Sela, and Gordon Wetzstein, as well as Florian Schroff and Hartwig Adam for continuous help in building this technology. We also thank Matthew Brown, Ricardo Martin-Brualla and Frederic Poitevin for their helpful feedback on the paper draft. We also acknowledge the use of the computational resources at the SLAC Shared Scientific Data Facility (SDF).

    AI / MLopen article
  • Google AI Blog15/03/2024, 18:22

    HEAL: A framework for health equity assessment of machine learning performance

    Posted by Mike Schaekermann, Research Scientist, Google Research, and Ivor Horn, Chief Health Equity Officer & Director, Google Core Health equity is a major societal concern worldwide with disparities having many causes. These sources include limitations in access to healthcare, differences in clinical treatment, and even fundamental differences in the diagnostic technology. In dermatology for example, skin cancer outcomes are worse for populations such as minorities, those with lower socioeconomic status, or individuals with limited healthcare access. While there is great promise in recent advances in machine learning (ML) and artificial intelligence (AI) to help improve healthcare, this transition from research to bedside must be accompanied by a careful understanding of whether and how they impact health equity. Health equity is defined by public health organizations as fairness of opportunity for everyone to be as healthy as possible. Importantly, equity may be different from equality. For example, people with greater barriers to improving their health may require more or different effort to experience this fair opportunity. Similarly, equity is not fairness as defined in the AI for healthcare literature. Whereas AI fairness often strives for equal performance of the AI technology across different patient populations, this does not center the goal of prioritizing performance with respect to pre-existing health disparities. Health equity considerations. An intervention (e.g., an ML-based tool, indicated in dark blue) promotes health equity if it helps reduce existing disparities in health outcomes (indicated in lighter blue). Health Equity Assessment of machine Learning performance (HEAL): a framework and dermatology AI model case study”, published in The Lancet eClinicalMedicine, we propose a methodology to quantitatively assess whether ML-based health technologies perform equitably. In other words, does the ML model perform well for those with the worst health outcomes for the condition(s) the model is meant to address? This goal anchors on the principle that health equity should prioritize and measure model performance with respect to disparate health outcomes, which may be due to a number of factors that include structural inequities (e.g., demographic, social, cultural, political, economic, environmental and geographic). The health equity framework (HEAL) Framework for Health Equity Assessment of machine Learning performance (HEAL). Our guiding principle is to avoid exacerbating health inequities, and these steps help us identify disparities and assess for inequitable model performance to move towards better outcomes for all. Case study on a dermatology model prior work. This example dermatology model was trained to classify 288 skin conditions using a development dataset of 29k cases. The input to the model consists of three photos of a skin concern along with demographic information and a brief structured medical history. The output consists of a ranked list of possible matching skin conditions. Using the HEAL framework, we evaluated this model by assessing whether it prioritized performance with respect to pre-existing health outcomes. The model was designed to predict possible dermatologic conditions (from a list of hundreds) based on photos of a skin concern and patient metadata. Evaluation of the model is done using a top-3 agreement metric, which quantifies how often the top 3 output conditions match the most likely condition as suggested by a dermatologist panel. The HEAL metric is computed via the anticorrelation of this top-3 agreement with health outcome rankings. We used a dataset of 5,420 teledermatology cases, enriched for diversity in age, sex and race/ethnicity, to retrospectively evaluate the model’s HEAL metric. The dataset consisted of “store-and-forward” cases from patients of 20 years or older from primary care providers in the USA and skin cancer clinics in Australia. Based on a review of the literature, we decided to explore race/ethnicity, sex and age as potential factors of inequity, and used sampling techniques to ensure that our evaluation dataset had sufficient representation of all race/ethnicity, sex and age groups. To quantify pre-existing health outcomes for each subgroup we relied on measurements from public databases endorsed by the World Health Organization, such as Years of Life Lost (YLLs) and Disability-Adjusted Life Years (DALYs; years of life lost plus years lived with disability). HEAL metric for all dermatologic conditions across race/ethnicity subpopulations, including health outcomes (YLLs per 100,000), model performance (top-3 agreement), and rankings for health outcomes and tool performance. (* Higher is better; measures the likelihood the model performs equitably with respect to the axes in this table.) HEAL metric for all dermatologic conditions across sexes, including health outcomes (DALYs per 100,000), model performance (top-3 agreement), and rankings for health outcomes and tool performance. (* As above.) HEAL metrics for all cancer and non-cancer dermatologic conditions across age groups, including health outcomes (DALYs per 100,000), model performance (top-3 agreement), and rankings for health outcomes and tool performance. (* As above.) Putting things in context Pareto condition (discussed further in the paper), which restricts model changes so that outcomes for each subpopulation are either unchanged or improved compared to the status quo, and performance does not worsen for any subpopulation. The HEAL framework, in its current form, assesses the likelihood that an ML-based model prioritizes performance for subpopulations with respect to pre-existing health disparities for specific subpopulations. This differs from the goal of understanding whether ML will reduce disparities in outcomes across subpopulations in reality. Specifically, modeling improvements in outcomes requires a causal understanding of steps in the care journey that happen both before and after use of any given model. Future research is needed to address this gap. Conclusion Acknowledgements The research described here is joint work across many teams at Google. We are grateful to all our co-authors: Terry Spitz, Malcolm Pyles, Heather Cole-Lewis, Ellery Wulczyn, Stephen R. Pfohl, Donald Martin, Jr., Ronnachai Jaroensri, Geoff Keeling, Yuan Liu, Stephanie Farquhar, Qinghan Xue, Jenna Lester, Cían Hughes, Patricia Strachan, Fraser Tan, Peggy Bui, Craig H. Mermel, Lily H. Peng, Yossi Matias, Greg S. Corrado, Dale R. Webster, Sunny Virmani, Christopher Semturs, Yun Liu, and Po-Hsuan Cameron Chen. We also thank Lauren Winer, Sami Lachgar, Ting-An Lin, Aaron Loh, Morgan Du, Jenny Rizk, Renee Wong, Ashley Carrick, Preeti Singh, Annisah Um'rani, Jessica Schrouff, Alexander Brown, and Anna Iurchenko for their support of this project.

    AI / MLopen article
  • Google AI Blog14/03/2024, 19:38

    Cappy: Outperforming and boosting large multi-task language models with a small scorer

    Posted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-following framework. This paradigm is exemplified by recent multi-task LLMs, such as T0, FLAN, and OPT-IML. First, multi-task data is gathered with each task following a task-specific template, where each labeled example is converted into an instruction (e.g., "Put the concepts together to form a sentence: ski, mountain, skier”) paired with a corresponding response (e.g., "Skier skis down the mountain"). These instruction-response pairs are used to train the LLM, resulting in a conditional generation model that takes an instruction as input and generates a response. Moreover, multi-task LLMs have exhibited remarkable task-wise generalization capabilities as they can address unseen tasks by understanding and solving brand-new instructions. The demonstration of the instruction-following pre-training of multi-task LLMs, e.g., FLAN. Pre-training tasks under this paradigm improves the performance for unseen tasks. FLAN-11B, T0-11B and OPT-IML-175B). As a result, operating such sizable models poses significant challenges because they demand considerable computational power and impose substantial requirements on the memory capacities of GPUs and TPUs, making their training and inference expensive and inefficient. Extensive storage is required to maintain a unique LLM copy for each downstream task. Moreover, the most powerful multi-task LLMs (e.g., FLAN-PaLM-540B) are closed-sourced, making them impossible to be adapted. However, in practical applications, harnessing a single multi-task LLM to manage all conceivable tasks in a zero-shot manner remains difficult, particularly when dealing with complex tasks, personalized tasks and those that cannot be succinctly defined using instructions. On the other hand, the size of downstream training data is usually insufficient to train a model well without incorporating rich prior knowledge. Hence, it is long desired to adapt LLMs with downstream supervision while bypassing storage, memory, and access issues. Certain parameter-efficient tuning strategies, including prompt tuning and adapters, substantially diminish storage requirements, but they still perform back-propagation through LLM parameters during the tuning process, thereby keeping their memory demands high. Additionally, some in-context learning techniques circumvent parameter tuning by integrating a limited number of supervised examples into the instruction. However, these techniques are constrained by the model's maximum input length, which permits only a few samples to guide task resolution. In “Cappy: Outperforming and Boosting Large Multi-Task LMs with a Small Scorer”, presented at NeurIPS 2023, we propose a novel approach that enhances the performance and efficiency of multi-task LLMs. We introduce a lightweight pre-trained scorer, Cappy, based on continual pre-training on top of RoBERTa with merely 360 million parameters. Cappy takes in an instruction and a candidate response as input, and produces a score between 0 and 1, indicating an estimated correctness of the response with respect to the instruction. Cappy functions either independently on classification tasks or serves as an auxiliary component for LLMs, boosting their performance. Moreover, Cappy efficiently enables downstream supervision without requiring any finetuning, which avoids the need for back-propagation through LLM parameters and reduces memory requirements. Finally, adaptation with Cappy doesn’t require access to LLM parameters as it is compatible with closed-source multi-task LLMs, such as those only accessible via WebAPIs. Cappy takes an instruction and response pair as input and outputs a score ranging from 0 to 1, indicating an estimation of the correctness of the response with respect to the instruction. Pre-training PromptSource that were used to train T0. This collection encompasses a wide range of task types, such as question answering, sentiment analysis, and summarization. Each dataset is associated with one or more templates that convert each instance from the original datasets into an instruction paired with its ground truth response. Cappy's regression modeling requires each pre-training data instance to include an instruction-response pair along with a correctness annotation for the response, so we produce a dataset with correctness annotations that range from 0 to 1. For every instance within a generation task, we leverage an existing multi-task LLM to generate multiple responses by sampling, conditioned on the given instruction. Subsequently, we assign an annotation to the pair formed by the instruction and every response, using the similarity between the response and the ground truth response of the instance. Specifically, we employ Rouge-L, a commonly-used metric for measuring overall multi-task performance that has demonstrated a strong alignment with human evaluation, to calculate this similarity as a form of weak supervision. As a result, we obtain an effective regression dataset of 160 million instances paired with correctness score annotations. The final Cappy model is the result of continuous pre-training using the regression dataset on top of the RoBERTa model. The pre-training of Cappy is conducted on Google's TPU-v4, with RedCoast, a lightweight toolkit for automating distributed training. Data augmentation with a multi-task LLM to construct a weakly supervised regression dataset for Cappy’s pre-training and fine-tuning. Applying Cappy Adapting multi-task LLMs with Cappy Downstream adaptation comparison between Cappy and approaches that rely on an LLM’s parameters, such as fine-tuning and prompt tuning. Cappy’s application enhances multi-task LLMs. Results PromptSource. We demonstrate that Cappy, with 360M parameters, outperforms OPT-175B and OPT-IML-30B, and matches the accuracy of the best existing multi-task LLMs (T0-11B and OPT-IML-175B). These findings highlight Cappy’s capabilities and parameter efficiency, which can be credited to its scoring-based pre-training strategy that integrates contrastive information by differentiating between high-quality and low-quality responses. On the contrary, previous multi-task LLMs depend exclusively on teacher-forcing training that utilizes only the ground truth responses. The overall accuracy averaged over eleven test tasks from PromptSource. “RM” refers to a pre-trained RLHF reward model. Cappy matches the best ones among existing multi-task LLMs. BIG-Bench, a set of manually curated tasks that are considered beyond the capability of many LLMs. We focus on all the 45 generation BIG-Bench tasks, specifically those that do not offer pre-established answer choices. We evaluate the performance using the Rouge-L score (representing the overall similarity between model generations and corresponding ground truths) on every test set, reporting the average score across 45 tests. In this experiment, all variants of FLAN-T5 serve as the backbone LLMs, and the foundational FLAN-T5 models are frozen. These results, shown below, suggest that Cappy enhances the performance of FLAN-T5 models by a large margin, consistently outperforming the most effective baseline achieved through sample selection using self-scoring of the LLM itself. The averaged Rouge-L score over 45 complex tasks within BIG-Bench. The x-axis refers to FLAN-T5 models of different sizes. Every dashed line represents an approach working on FLAN-T5s. Self-scoring refers to using the cross-entropy of LLM to select responses. Cappy enhances the performance of FLAN-T5 models by a large margin. Conclusion Acknowledgments Thanks to Bowen Tan, Jindong Chen, Lei Meng, Abhanshu Sharma and Ewa Dominowska for their valuable feedback. We would also like to thank Eric Xing and Zhiting Hu for their suggestions.

    AI / MLopen article
  • Google AI Blog12/03/2024, 21:15

    Talk like a graph: Encoding graphs for large language models

    Posted by Bahare Fatemi and Bryan Perozzi, Research Scientists, Google Research Imagine all the things around you — your friends, tools in your kitchen, or even the parts of your bike. They are all connected in different ways. In computer science, the term graph is used to describe connections between objects. Graphs consist of nodes (the objects themselves) and edges (connections between two nodes, indicating a relationship between them). Graphs are everywhere now. The internet itself is a giant graph of websites linked together. Even the knowledge search engines use is organized in a graph-like way. Furthermore, consider the remarkable advancements in artificial intelligence — such as chatbots that can write stories in seconds, and even software that can interpret medical reports. This exciting progress is largely thanks to large language models (LLMs). New LLM technology is constantly being developed for different uses. Since graphs are everywhere and LLM technology is on the rise, in “Talk like a Graph: Encoding Graphs for Large Language Models”, presented at ICLR 2024, we present a way to teach powerful LLMs how to better reason with graph information. Graphs are a useful way to organize information, but LLMs are mostly trained on regular text. The objective is to test different techniques to see what works best and gain practical insights. Translating graphs into text that LLMs can understand is a remarkably complex task. The difficulty stems from the inherent complexity of graph structures with multiple nodes and the intricate web of edges that connect them. Our work studies how to take a graph and translate it into a format that an LLM can understand. We also design a benchmark called GraphQA to study different approaches on different graph reasoning problems and show how to phrase a graph-related problem in a way that enables the LLM to solve the graph problem. We show that LLM performance on graph reasoning tasks varies on three fundamental levels: 1) the graph encoding method, 2) the nature of the graph task itself, and 3) interestingly, the very structure of the graph considered. These findings give us clues on how to best represent graphs for LLMs. Picking the right method can make the LLM up to 60% better at graph tasks! Pictured, the process of encoding a graph as text using two different approaches and feeding the text and a question about the graph to the LLM. Graphs as text GraphQA. Think of GraphQA as an exam designed to evaluate powerful LLMs on graph-specific problems. We want to see how well LLMs can understand and solve problems that involve graphs in different setups. To create a comprehensive and realistic exam for LLMs, we don’t just use one type of graph, we use a mix of graphs ensuring breadth in the number of connections. This is mainly because different graph types make solving such problems easier or harder. This way, GraphQA can help expose biases in how an LLM thinks about the graphs, and the whole exam gets closer to a realistic setup that LLMs might encounter in the real world. Overview of our framework for reasoning with graphs using LLMs. Erdős-Rényi, scale-free networks, Barabasi-Albert model, and stochastic block model, as well as simpler graph structures like paths, complete graphs, and star graphs, providing a diverse set of data for training. When working with graphs, we also need to find ways to ask graph-related questions that LLMs can understand. Prompting heuristics are different strategies for doing this. Let's break down the common ones: Zero-shot: simply describe the task ("Is there a cycle in this graph?") and tell the LLM to go for it. No examples provided. Few-shot: This is like giving the LLM a mini practice test before the real deal. We provide a few example graph questions and their correct answers. Chain-of-Thought: Here, we show the LLM how to break down a problem step-by-step with examples. The goal is to teach it to generate its own "thought process" when faced with new graphs. Zero-CoT: Similar to CoT, but instead of training examples, we give the LLM a simple prompt, like "Let's think step-by-step," to trigger its own problem-solving breakdown. BAG (build a graph): This is specifically for graph tasks. We add the phrase "Let's build a graph..." to the description, helping the LLM focus on the graph structure. We explored different ways to translate graphs into text that LLMs can work with. Our key questions were: Node encoding: How do we represent individual nodes? Options tested include simple integers, common names (people, characters), and letters. Edge encoding: How do we describe the relationships between nodes? Methods involved parenthesis notation, phrases like "are friends", and symbolic representations like arrows. Various node and edge encodings were combined systematically. This led to functions like the ones in the following figure: Examples of graph encoding functions used to encode graphs via text. Analysis and results How LLMs handle graph tasks LLMs struggle: On most of these basic tasks, LLMs did not do much better than a random guess. Encoding matters significantly: How we represent the graph as text has a great effect on LLM performance. The "incident" encoding excelled for most of the tasks in general. Our results are summarized in the following chart. Comparison of various graph encoder functions based on their accuracy on different graph tasks. The main conclusion from this figure is that the graph encoding functions matter significantly. Bigger is (usually) better PaLM 2. Here is a summary of our findings: In general, bigger models did better on graph reasoning tasks. It seems like the extra parameters gave them space to learn more complex patterns. Oddly, size didn't matter as much for the “edge existence” task (finding out if two nodes in a graph are connected). Even the biggest LLM couldn't consistently beat a simple baseline solution on the cycle check problem (finding out if a graph contains a cycle or not). This shows LLMs still have room to improve with certain graph tasks. Effect of model capacity on graph reasoning task for PaLM 2-XXS, XS, S, and L. Do different graph shapes confuse LLMs Samples of graphs generated with different graph generators from GraphQA. ER, BA, SBM, and SFN refers to Erdős–Rényi, Barabási–Albert, Stochastic Block Model, and Scale-Free Network respectively. Comparing different graph generators on different graph tasks. The main observation here is that graph structure has a significant impact on the LLM’s performance. ER, BA, SBM, and SFN refers to Erdős–Rényi, Barabási–Albert, Stochastic Block Model, and Scale-Free Network respectively. Conclusion How to translate the graph to text: how we represent the graph as text significantly influences LLM performance. The incident encoding excelled for most of the tasks in general.. Task type: Certain types of graph questions tend to be harder for LLMs, even with a good translation from graph to text. Graph structure: Surprisingly, the "shape" of the graph that on which we do inference (dense with connections, sparse, etc.) influences how well an LLM does. This study revealed key insights about how to prepare graphs for LLMs. The right encoding techniques can significantly boost an LLM's accuracy on graph problems (ranging from around 5% to over 60% improvement). Our new benchmark, GraphQA, will help drive further research in this area. Acknowledgements We would like to express our gratitude to our co-author, Jonathan Halcrow, for his valuable contributions to this work. We express our sincere gratitude to Anton Tsitsulin, Dustin Zelle, Silvio Lattanzi, Vahab Mirrokni, and the entire graph mining team at Google Research, for their insightful comments, thorough proofreading, and constructive feedback which greatly enhanced the quality of our work. We would also like to extend special thanks to Tom Small for creating the animation used in this post.

    AI / MLopen article