The reading room
Engineering blogs
worth your time.
Recent posts and AI research from the companies I follow, pulled from their RSS feeds and linked to the originals. Filter by company, newest first.
1189 posts tracked
SalesforceHow Standardizing Product Telemetry Reduced Time to Insight by 97%By Rounak Mehta, Adrian Eng, Sandeep Singh, and Kishore Kanchapalli. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Adrian Eng, Dire…OpenAIVirgin Atlantic sharpens customer journeys with ChatGPT WorkVirgin Atlantic is accelerating research, product planning, and decision-making with ChatGPT Work, helping teams connect signals across the customer journey.OpenAIHow Zapier transformed core marketing processes with ChatGPT WorkThe enterprise marketing team at Zapier uses ChatGPT Work to reduce the number of drop-offs in its lead funnel, build campaign assets, and automate reporting.OpenAIPremium seats are coming to ChatGPT BusinessPremium seats are coming to ChatGPT Business. Sign up by August 20 to get $100 in workspace credits and unlock higher usage for your team's most demanding work.OpenAIPutting frontier cyber models in more trusted handsApproved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.OpenAIExpanding Daybreak as the Cyber Defense Window NarrowsMeet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit validation, and security testing.OpenAIWhat building an AI-native finance function taught meOpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and AI ROI.OpenAIModel ML completes finance work more efficiently with GPT-5.6 SolModel ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.OpenAIOpenAI’s letter to Governor Abbott on responsible AI infrastructure in TexasOpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. The letter supports reliable, transparent growth that benefits Texans.NVIDIARun Local Agentic AI Workflows with Meta’s Muse Glimmer on NVIDIAMeta returns to the open source ecosystem with the release of Muse Glimmer, a 30B open-weight dense model with a 120K+ context window built for local AI...Hugging FaceMeta is back with Muse Glimmer: local, agentic, multimodal, and open sourceHugging FaceMaking Knowledge Distillation Cheap Enough to Run at ScaleHugging FaceBuild Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTSGitHubUsing the GitHub Copilot SDK for JavaEnterprise Java developers have a new superpower—drive GitHub Copilot from idiomatic Java code with annotations, virtual threads, and more. The post Using the GitHub Copilot SDK for Java appeared first on The GitHub Blog…CloudflareServing the most critical missions: Cloudflare for Government achieves FedRAMP Class D (High) Certified statusCloudflare for Government achieves FedRAMP Class D (High) Certified status. We also announce our commitment to pursue DoD IL4 authorization. Cloudflare brings world-class security, performance, and developer products to…CloudflareEverything we launched during Agents WeekOur latest Agents Week has come to a close. Here’s a recap of all the announcements we made, from Wallets to Radar.AWSAWS Weekly Roundup: AWS Heroes Summit, Web Search on Amazon Bedrock, Dogwood, Kiro Crew, and more (August 10, 2026)Last week, we brought together AWS Heroes from around the world to connect, collaborate, and celebrate the builders who go above and beyond for the AWS community. The AWS Heroes Summit, an invite-only annual gathering, b…arXiv · NLPFailure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation SystemA long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominated by source or prompt material, or interrupted after producing text…arXiv · NLPEmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language ModelsDespite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception,…arXiv · NLPUNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text ClassifiersNeural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on…arXiv · NLPReading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision ModelInverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produce…arXiv · NLPMoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask ExpertsLarge language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptati…arXiv · NLPBusiness Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B BaselineLLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema c…arXiv · NLPVerifiably grounded machine interpretation of lunar geologyPlanetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an automated "machine intelligence geologist" by embedding this distinct…arXiv · NLPIs the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Check…arXiv · NLPComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with CouponsReal-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeou…arXiv · NLPAccurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL WritingSecond language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipelin…arXiv · NLPBeyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM AgentsSelf-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability…arXiv · NLPUniversal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic LanguagesClosely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare two approaches to script unification: the general-purpose uroman ro…arXiv · NLPTemporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax LawWe identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treat…arXiv · NLPIntent Speaks Louder: Controllable User Simulation Beyond Response ImitationUser simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support mult…arXiv · NLPReducing Pretraining-Generation Mismatch in Diffusion Language ModelsAutoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native…arXiv · NLPZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language ModelsTransformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitati…arXiv · NLPDepth-adaptive Inference of Looped Language Models via Continuous Depth BatchingA main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. How…arXiv · NLPLearning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement LearningNatural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplyi…arXiv · NLPBuild it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media postsDetecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measu…arXiv · NLPTCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research AbilityWe introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at…arXiv · NLPMawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance DetectionPublicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-v2 Extension, consisting of 996 manually ann…arXiv · NLPELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language ModelsLarge language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is…arXiv · NLPPragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language ModelsIn the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks direc…arXiv · NLPSe-DPO: Self-Evolving Token Credit for Direct Preference OptimizationDirect Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of indi…arXiv · NLPMDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQLTraditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in…arXiv · NLPMeasuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful JailbreaksInternal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also cat…arXiv · NLPAvalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game MechanicsTheory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.…arXiv · NLPListwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure RerankingReranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial lexical gap between patient language and clinical nomenclature. We pres…arXiv · NLPVeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative ScaffoldingGreat fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot yet smell of decay, weaving domain expertise into the fabric of invente…arXiv · NLPMatryoshka Language Model SuitesTraining a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a singl…arXiv · NLPHow Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and HumansLarge language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social a…arXiv · NLPREFRAMED: Towards Realistic Audio Description Generation for MoviesAudio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be ins…arXiv · NLPCultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation RobustnessMultilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlook…arXiv · NLPStructured Phonological Representations for Audio-Articulatory rtMRI Speech ClassificationReal-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an au…arXiv · NLPPragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language ModelsLarge Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial…arXiv · NLPKGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMsAnswering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to…arXiv · NLPComparing British and American Audio Description of MoviesNarrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than…arXiv · NLPSWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code RefactoringAs AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found tha…arXiv · NLPParameter Exploration for RLVR via Variational LearningExploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significan…arXiv · NLPMacaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRAMacaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursu…arXiv · NLPRA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classificationFinancial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existing work on resource-efficient financial NLP has largely focused on c…arXiv · NLPMismatch Matters: On-Policy Distillation Beyond Token AgreementOn-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect toke…arXiv · NLPAgentic Auto-Research is Fuzz TestingAutonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue…arXiv · NLPTowards Expert-level Medical AI for Real-time Video ConsultationsAudio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards…arXiv · NLPFusion Training for Mathematical Generalization in Large Language ModelsThinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics…arXiv · NLPConsilience for Verifier-Free Test-Time ScalingTest-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-T…arXiv · NLPDecoding-Level Taboo: A Diagnostic Stress Test for LLM RobustnessLarge language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world dep…arXiv · NLPFrom Values to Benchmarks: Evaluating Large Language Models for Governmental Use in DutchLarge language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English co…arXiv · NLPMultimodal Model Diffing for Feature Discovery and ControlMultimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection,…arXiv · NLPBeyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded DimensionsAutomated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture…arXiv · MLFlow-based conditional cardiac anatomy generation for virtual cohortsCardiac digital twin research is moving from subject-specific anatomical replicas toward virtual cohorts that represent clinically relevant population subgroups. Yet access to representative imaging-derived anatomy datas…arXiv · MLMixFormer: Linear Transformer with Mixture of Memory ExpertsState Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in long-context modeling. However, existing SSMs suffer from limited input…arXiv · MLHierarchical rank-evolving representation for physics-informed neural networksRecently, tensor-based physics-informed neural networks (T-PINNs) have received increasing attention. However, existing T-PINNs still face a fundamental challenge: they mainly rely on pre-specified low-rank tensor decomp…arXiv · MLWhen Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space CompositionTask arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predictable changes in model function. We separate parameter geometry fro…arXiv · MLWalk-on-Spheres Monte Carlo and deep neural network approximations of elliptic PDEs with drift and killingIn this paper we provide Monte Carlo and deep neural network approximations for stochastic representations of solutions to linear elliptic partial differential equations with constant diffusion, drift and killing. Buildi…arXiv · MLTracking the Best Strategy in an Extensive-Form GameWe consider the extensive-form bandit problem where on each trial the learner plays an extensive-form game against an oblivious adversary. We focus on the notion of switching regret, which measures the expected performan…arXiv · MLXFeat Revisited: Reproducibility and Evaluation of a Lightweight Image MatcherWe present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the arc…arXiv · MLGeneralized Convexity and Smoothness via Conjugate Duality: Optimization Theory for Deep Neural NetworksDeep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success. This limitation ari…arXiv · MLDual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMsLarge reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direct refusals or safety rationales, yet oft…arXiv · MLTraining-Free Universal Approximation by Prompting Random TransformersHow expressive is prompting a transformer? Answering this question is important for separating the roles of prompting, architecture, and pretraining in transformer models, and for determining whether task-specific behavi…arXiv · MLDistributed Optimization with Streaming Data: A Temporal Weighting PerspectiveOptimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant objective functions, many modern applications operate in dynamic environments where…arXiv · MLHyperbolic Multimodal Continual LearningHyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities. Despite this progress, how such representat…arXiv · MLLEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNNGraph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable with depth, and over-squashing, where long-range information is compressed through…arXiv · MLStructure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane DetectionLane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based detectors provide efficient candidate generation, their performance is…arXiv · MLBayesian Symbolic Regression with Entropic Reinforcement LearningSymbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model stru…arXiv · MLSatellite Trajectory Optimization via Proximal Policy Optimization for Space Debris AvoidanceCollision avoidance systems are commonly used to avoid fragmentation events occurring in Low-Earth Orbit (LEO) and Geosynchronous Equatorial Orbit (GEO). However, these events have been growing in frequency as orbital co…arXiv · MLLoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack DetectionFace presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-datas…arXiv · MLActivation Probes Surface Code-Security Signals that the Model's Output MissesAI coding agents now write a growing share of production code, and human security review does not scale at the rate code is generated. The agents in widest use are closed-weight, so a deploying team cannot read their int…arXiv · MLDeep Learning Imputation of Missing Radius of Maximum Winds (Rmax) Values in Tropical Cyclone Best-Track DataProbabilistic coastal hazard assessments require accurate characterization of tropical cyclone (TC) parameters, yet datasets often contain missing records for the radius of maximum winds (Rmax), a key variable in Joint P…arXiv · MLFedOrbit: Adaptive Personalized Federated Learning for Non-IID LEO Satellite ConstellationsFederated learning (FL) in Low Earth Orbit (LEO) satellite constellations is affected by non-IID data and irregular ground-station visibility, both driven by orbital geometry. Global aggregation performs poorly when orbi…arXiv · MLConfusion-Geometry Rebalancing for Long-Tailed Adversarial TrainingAdversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner maximization may further amplify this bi…arXiv · MLRecurrent Neural Networks Beyond Time: Learning from Multiple Ordered ProjectionsRecurrent neural networks (RNNs) are widely used for sequence learning, yet their application is commonly associated with temporal data, although recurrent computation fundamentally operates on ordered sequences rather t…arXiv · MLEvaluating Generative Time-Series Models on Data with Point MassesMany of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such d…arXiv · MLTest-Time Scaling for CAD Generation via Verifier-Free Consensus SelectionLarge language models can write parametric CAD programs from a natural-language description (text-to-CAD generation), but a single sample is often wrong. Increasing test-time compute by sampling multiple candidates only…arXiv · MLInput convex neural networks as surrogates in mathematical optimisationEmbedding trained neural networks as surrogates within optimisation problems is an established practice in operations research. The prevailing approach uses feedforward neural networks (FNNs) with ReLU activations, whose…arXiv · MLPET/CT Radiogenomic Mutation Prediction in Non-Small Cell Lung Cancer Using Multi-Label LearningLung cancer remains one of the leading causes of cancer- related mortality worldwide. Although targeted therapies have improved outcomes for patients with non-small cell lung cancer (NSCLC), they rely on mutation profili…arXiv · MLRethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive ApproachLow-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired b…arXiv · MLSR-OPSD: Self-Referenced On-Policy Self-DistillationOn-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome…arXiv · MLDefining Decentralization: An Ontological PerspectiveDecentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, cloud infrastructures,…arXiv · MLDisentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert RoutingRetinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution. We…arXiv · MLMoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General GeometriesTransformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in l…arXiv · MLReliableNet: A Chance-Constrained Approach to Trustworthy Classification in Deep LearningA prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken. Empirical risk minimization (ERM) controls average l…arXiv · MLC$^2$A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray ClassificationThoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding \emph{where} findings lie and \emph{how} they co-occur. We propose \textbf{C$\mathbf{^2}$…arXiv · MLAirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality ForecastingAccurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentrati…arXiv · MLParameter Exploration for RLVR via Variational LearningExploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significan…arXiv · MLMacaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRAMacaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursu…arXiv · MLDistill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-DistillationReinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Sk…arXiv · MLMulti-Agent AI Safety as an Institutional Design ProblemAI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior.…arXiv · MLDeep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive BehaviorsBody-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and anxiety disorders. Their early, objective detection remains difficu…arXiv · MLRA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classificationFinancial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existing work on resource-efficient financial NLP has largely focused on c…arXiv · MLReal-Time Climate Risk Assessment for Supply Chain Resilience: A Data-Driven Nowcasting Framework for Colombian AgricultureThis paper presents a methodological framework for real-time climate risk assessment using data-driven nowcasting techniques to enhance supply chain resilience in Colombian agricultural contexts. Climate variability in C…arXiv · MLRynnValue: Scaling Robotic Value Foundation Models with Temporal DistanceGeneral-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing appr…arXiv · MLStealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side,…arXiv · MLLogarithmic-Free Moment and Generalization Bounds for Uniformly Stable AlgorithmsUniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum…arXiv · MLFinancial Numerical Prediction and Allocation as Token GenerationFinancial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ultimately evaluated. We investigate whether a causal language model ca…arXiv · MLSpace-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast FootballBall possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP…arXiv · MLBDH-CQ: In-Context Learning with Recurrent Latent ReasoningWe introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query…arXiv · MLConsilience for Verifier-Free Test-Time ScalingTest-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-T…arXiv · MLFairness in Link Prediction Beyond Demographic Parity: A Reproducibility StudyIn fair ranked link prediction, demographic parity ($Δ_\mathrm{DP}$) is a common fairness metric. Yet, Mattos et al. (2025) argue that it fails to detect exposure bias because it ignores where links appear in the ranking…arXiv · MLMultimodal Model Diffing for Feature Discovery and ControlMultimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection,…arXiv · AILEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNNGraph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable with depth, and over-squashing, where long-range information is compressed through…arXiv · AITSPORec: Token Selection via Preference Optimization for LLM-Based Sequential RecommendationLarge Language Models (LLMs) have emerged as powerful tools for improving recommendation systems. The effectiveness of LLMs arises from their ability to harness rich textual information and their capacity to model hetero…arXiv · AIStructure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane DetectionLane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based detectors provide efficient candidate generation, their performance is…arXiv · AIAdaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree SearchReliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives against multiple competing failure mechanisms. Current practice relies on static t…arXiv · AIMeasuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful JailbreaksInternal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also cat…arXiv · AIRethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific…arXiv · AINeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron SegmentationAccurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses significant challenges to existing segmentation methods. These method…arXiv · AIDUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video GenerationDiffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--…arXiv · AIAvalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game MechanicsTheory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight.…arXiv · AIPredictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controllerRecent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since learning-based control may not be able to realize safety-guaranties, i…arXiv · AIHallucination-Free GUI Grounding via Regression-Free Layout-Aware MatchingGUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abs…arXiv · AIOpen Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative ModelsRecent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hundreds or thousands of images or videos, which is computationally exp…arXiv · AIAdaptive Semantic Capacity Allocation for Parallel Generative RecommendationAutoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identifiers. Parallel generation methods alleviate this bottleneck by predicting all se…arXiv · AIConfusion-Geometry Rebalancing for Long-Tailed Adversarial TrainingAdversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner maximization may further amplify this bi…arXiv · AIEvaluating Generative Time-Series Models on Data with Point MassesMany of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no ride is requested, no part is ordered. We report what happens when such d…arXiv · AIModel Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world modelsPredicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, causal model, not a curve fit; and learning such a model requires \emph{experiment…arXiv · AIMatryoshka Language Model SuitesTraining a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a singl…arXiv · AIHow Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and HumansLarge language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social a…arXiv · AIColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill ScannersAgent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skills, leaving risks fr…arXiv · AIRethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive ApproachLow-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an efficient way to fine-tune large models in federated learning paradigm. Inspired b…arXiv · AISR-OPSD: Self-Referenced On-Policy Self-DistillationOn-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome…arXiv · AIDefining Decentralization: An Ontological PerspectiveDecentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, cloud infrastructures,…arXiv · AISecond-Order Muon Done Right: A Principled Marriage of Spectral Geometry and CurvatureMuon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and reuses it across several optimization steps. Conditioned on any positive-definite l…arXiv · AIMoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General GeometriesTransformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observations into compact latent tokens and learning physical interactions in l…arXiv · AICultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation RobustnessMultilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlook…arXiv · AIAirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality ForecastingAccurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because pollutant channels differ in periodicity and distribution drift, while their concentrati…arXiv · AIKGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMsAnswering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to…arXiv · AICARD: Controlled Agentic Reddit Discussions for Credit Card SimulationOnline credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated…arXiv · AIModern Backbones Improve Multi-task DETR for Mammography Classification and Lesion LocalizationJoint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both im…arXiv · AIParameter Exploration for RLVR via Variational LearningExploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significan…arXiv · AIMedPixel: A Unified Pixel-Language Model for Medical Reasoning and SegmentationReliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segme…arXiv · AIDistill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-DistillationReinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong, which account for 63.0-68.0% of groups in our experiments. We propose SKALD (Sk…arXiv · AIMulti-Agent AI Safety as an Institutional Design ProblemAI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior.…arXiv · AIMismatch Matters: On-Policy Distillation Beyond Token AgreementOn-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect toke…arXiv · AICEAA: A Cognitive Embodied Agents Architecture for Interactive Computing SystemsThe development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing arc…arXiv · AIAgentic Auto-Research is Fuzz TestingAutonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue…arXiv · AIAgentic Harnesses: LLM-Driven Verification Layers for Robot AutonomyAdvances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning mode…arXiv · AITowards Expert-level Medical AI for Real-time Video ConsultationsAudio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards…arXiv · AIStealing Reasoning Traces from Proprietary LLM APIsLeading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side,…arXiv · AISci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science DomainsWe introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across fou…arXiv · AIArchAgent v2: A Case Study with the Data Prefetching ChampionshipAgentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search spaces, strict hardwar…arXiv · AIEnergy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion PlanningPhysically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to real-world execution dynamics. While latent world models offer a promising approach…arXiv · AISHE: Trajectory-driven Safety Harness Evolution for LLM AgentsThe safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often trea…arXiv · AIBDH-CQ: In-Context Learning with Recurrent Latent ReasoningWe introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuously update the model's recurrent memory; the model then solves a query…arXiv · AIFusion Training for Mathematical Generalization in Large Language ModelsThinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics…arXiv · AIDSLE: A Learning Environment for Dark Souls Boss EncountersWe introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-style interface. DSLE…arXiv · AIGENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid AnalysisFoundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power system analysis, where strict physical consistency must be enforced.…arXiv · AIFrom Values to Benchmarks: Evaluating Large Language Models for Governmental Use in DutchLarge language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English co…arXiv · AIMultimodal Model Diffing for Feature Discovery and ControlMultimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection,…arXiv · AIBeyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded DimensionsAutomated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture…OpenAIHow HSP GRUPPE builds AI capabilities for tax advisoryDiscover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and client service.OpenAIResponding to the next frontier of critical cyber capabilitiesOpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.NetflixHow and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC…How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC execution API Authors: Nilesh Mishra and Ajit Koti This is the third entry of a multi-part blog series describing how we buil…Hugging FaceTutorMoments: Do AI tutors know when to help and when to hold back?CloudflareUnifying Workers AI and AI Gateway into a single AI control planeCloudflare is unifying AI Gateway and Workers AI into a single control plane, giving developers observability, billing, and dynamic routing across both managed GPUs and external providers. Learn how unified bindings and…CloudflareAnnouncing Cloudflare Ambassadors, Community Engineers, and another $1M in open-source fundingWe are launching updated community programs, including Cloudflare Ambassadors and Community Engineers, backed by $1M in open-source funding. Learn how we are supporting maintainers and scaling our developer community.CloudflareIntroducing Radar Researcher: An AI tool for exploring Internet data in plain languageCloudflare Radar Researcher is a new AI-powered tool that lets you explore global Internet trends and traffic data using plain language. Built entirely on Cloudflare's Developer Platform, it turns natural language querie…CloudflareUnveiling good and bad behaviors on the Agentic InternetCloudflare is shifting bot mitigation from point-in-time Risk assessment to continuous Trust evaluation. Learn how new good and bad behaviors from bots and agents are assessed by our systems, including BotBase and Precur…AppleArbitrage: Efficient Reasoning via Advantage-Aware SpeculationModern Large Language Models achieve impressive reasoning capabilities with long Chain of Thoughts, but they incur substantial computational cost during inference, and this motivates techniques to improve the performance…AppleBeyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language ModelsLarge Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive Language Models (ARMs…AppleScaling Categorical Flow MapsContinuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalit…AnthropicImproving Fable 5's biology safeguardsOpenAIFrom asking to doing: How the world is putting ChatGPT to workNew OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.OpenAIWorking with the American Psychological Association on youth mental health and AIOpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health.OpenAIImproving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free usersChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna.Hugging FaceBaseten on Hugging Face Inference Providers 🔥GitLabConfidential AI for GitLab Self-HostedYour developers want AI coding agents. Your source code is regulated IP that can t be sent to a third-party AI service, and your compliance team has said so in writing. The usual escape hatch, standing up your own GPU cl…GitLabGitLab Secrets Manager adds ESO, Terraform, API supportToday, you might maintain separate secret stores for CI/CD, Kubernetes, and Terraform. However, that leaves multiple tools to manage, access models to keep in sync, and audit trails to correlate when something goes wrong…GitHubHow we took malware advisories beyond npmGitHub malware advisories no longer stop at npm. Here's how we wired OpenSSF's malicious-packages data into the Advisory Database, and why we built the pipeline paranoid. The post How we took malware advisories beyond np…GitHubA guide to slash commands in the GitHub Copilot appGo beyond chat in the GitHub Copilot app with these slash commands. They'll help you plan, collaborate, automate, and customize your dev workflow. The post A guide to slash commands in the GitHub Copilot app appeared fir…DiscordGeneral Availability of Mobile Platform Support Arrives in Discord Social SDK Version 1.10Mobile support is now generally available in the Discord Social SDK, allowing developers to extend Discord s social experiences to iOS and Android.DiscordDiscord Patch Notes: August 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.Google DeepMindWeatherNext: AI model achieves breakthrough in forecasting cyclonesCohereCohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipelineCloudflareGive any website a WebMCP interfaceToday we're launching a developer preview of WebMCP on Cloudflare. With one switch, any site becomes usable by browser AI agents — no new APIs, no origin changes — while the human stays in control and creators keep their…CloudflareIntroducing Kitesurf: The agent-first browser that runs in V8 isolates on Cloudflare WorkersWe should be giving all agents tools that excel at what’s important for an AI model. Kitesurf is Cloudflare’s new stateless, highly scalable, and cost-effective web browser that runs entirely on top of Workers and was de…CloudflareBuilding an open Agentic Internet: readable, discoverable, callable, and payableAgents are a new kind of visitor. They don't render CSS or click ads, but they have a paying human on the other end. Block them and you block your customer. We're building the open tools and protocols so publishers and a…CloudflareFrom ranking to recommended: get your site ready to thrive in the age of AI agentsMore than half of requests now come from machines, not people. Agent Readiness shows how well agents can discover and read your site, while Answer Engine Optimization tracks how often AI assistants recommend you.CloudflareThe next generation of MCPThe next version of MCP has a rewritten, stateless core that just works on Workers. We cover upgrades to the protocol, the new feature lifecycle and SDK migration path, and hear from early adopters already running it in…CloudflareCloudflare AI Search: give your agents a search engine for your dataAI Search makes search easier than ever, with no Cloudflare primitives to stitch together. Point it at your data to create a search for your own files and websites. We're also sharing a preview of our new pricing model.AWSRuntime instances: persistent compute for production AI agents on Amazon Bedrock AgentCoreAnnouncing runtime instances in Amazon Bedrock AgentCore—persistent, managed EC2 infrastructure for production AI agents with multi-agent collaboration, GPU support, and sessions lasting up to 14 days.AppleDeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer CompletenessLarge language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as “Which actor from…AppleLocking Pretrained Weights via Deep Low-Rank Residual DistillationThe quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also all…SalesforceHow Salesforce Eliminated Single-Region Risk and Reduced Downtime Blast Radius at 4B Metrics/MinBy Priyanka Singla, Aaron Stockton, and Vincent Poon In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priyanka Singla, Senior Software…MetaFrom User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads RankingEvery day, Meta’s recommendation platforms handle billions of user interactions, generating rich temporal signals that capture individual preferences and intent across products, ads, and content. In our 2024 post on sequ…CloudflareCatching rogue AI behavior with identity-aware analyticsIdentity-aware AI Gateway is now in open beta. User Insights turns that traffic into a behavioral baseline for every person and agent, and flags insider risk the moment it appears.CloudflareWriteGuard: fine-grained controls for MCP ServersAt Cloudflare, we knew we could not depend on every employee to configure every agent perfectly or watch every tool call. Before expanding write access across our own internal MCP servers, we built WriteGuard. We are now…CloudflareCloudflare OS: an open platform for agents, apps, and workCloudflare OS is an open-source platform that lets everyone in your company build apps, automate work, and safely access internal systems, shaped around what your organization knows and how it operatesCloudflareHow we’re rethinking work at Cloudflare with Cloudflare OSWe built Cloudflare OS to equip our teams to safely rethink how they get work done with AI. The platform brings together the best of our technologies, from our Compute primitives to our Zero Trust suite. This post walks…CloudflareThe Agent Access ModelThe Agent Access Model proposes a new architecture to secure task-scoped agents using strict identity brokering, continuous mediation, and stateful trust.CloudflareCloudflare is the only vendor named a Visionary in 2026 SASE and SSE reportsWe're honored to announce that Cloudflare is the only vendor that has been recognized as a Visionary in both the 2026 Gartner® Magic Quadrant™ for SASE Platforms and the 2026 Gartner® Magic Quadrant™ for Security Service…AWSAmazon DynamoDB now supports real-time vector search at any scaleDynamoDB now supports native vector search with single-digit millisecond latency at 99%+ recall. It is designed for any scale, even trillions of vectors and requires zero infrastructure management.AppleTaming Outlier Tokens in Diffusion TransformersWe study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention w…YelpMigrating a Large Flow Monorepo to TypeScriptIn early 2017, Webcore selected Flow as Yelp’s next-generation typechecker over TypeScript. At the time there was no clear frontrunner. Flow had better support for React, better performance, and a respectable repository…OpenAINew ways to learn and teach with ChatGPT Work and CodexExplore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.OpenAIThird-party cyber evaluations involving OpenAI modelsOpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.NVIDIAGenerate Trajectories, Reasoning Traces, and Auto-Labels with NVIDIA Alpamayo 2 SuperAutonomous vehicle (AV) development often relies on separate models for trajectory generation, high-level intent prediction, scene understanding, and data...NVIDIABeyond VLAs: How World Action Models Reshape Robot ManipulationA central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene...Mistral AIIntroducing Shieldstral.Hugging FaceDeploy local agents everywhere with LFM2.5-2.6BGitHubTurn one giant AI-generated pull request to a reviewable stackInstead of one huge, un-reviewable pull request, teach coding agents to decompose work into a clean, ordered stack with GitHub stacked pull requests. The post Turn one giant AI-generated pull request to a reviewable stac…GitHubHow the GitHub legal team used Copilot CLI to streamline their workflowsLearn how to build tools to simplify how you work—without writing a single line of code. The post How the GitHub legal team used Copilot CLI to streamline their workflows appeared first on The GitHub Blog .DiscordSquircles, Styles, and Spacing: How Your Feedback is Helping Improve MobileWe re working to more closely match your mobile and desktop experiences, including changes to server icons and the chat bar. (Did we mention Ash, the OG dark theme, is back?)CloudflareAnnouncing Cloudflare Wallets: the programmable wallet for the agentic InternetCloudflare Wallets will provide AI agents with native payments and verifiable identity on the web. Using the x402 protocol, agents can autonomously purchase APIs and content within clear safety guardrails.CloudflareThe Agent Development Lifecycle has arrived on CloudflareAgents can write code faster than teams can review, deploy, and maintain it. Today we’re introducing the Agent Development Lifecycle and the Cloudflare primitives that underpin itCloudflareHow we built a software factory to drive Astro’s GitHub issue count to zeroBy replacing manual issue verification with isolated AI subagents running in GitHub Actions, the Astro maintainers reduced open issue count by 85%. This post explores the architecture behind automated bug reproduction, p…CloudflareYour agent can now debug Workers with local tracingwrangler dev now produces structured traces for every local request. Your coding agent can hit a single API to pinpoint exactly what failed and why — no deployment required.CloudflareIntroducing: Cloudflare AgentsCloudflare Agents brings all of your deployed agent sessions into a single experience, surfacing key information and insights into how your agents perform at scale.CloudflareRun CI/CD for millions of repos — on your platform, on CloudflareLearn how to build customizable, sandboxed CI/CD pipelines natively on Cloudflare using Workflows, Artifacts, and the CI SDK. We walk through replacing complex YAML configurations with TypeScript workflow steps and self-…CloudflareHow Cloudflare enforces engineering standards using AIWe created the Cloudflare Codex, a governed body of engineering standards that AI agents consume across the development lifecycle. By pairing structured RFCs with agentic reviews, teams automatically enforce consistency…AnthropicMariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs OfficerSalesforceRemoving the Security Barrier to Agentforce AdoptionBy Pallavi Rajan Udmalpet and Janhavi Deshpande. As enterprises accelerate secure AI adoption with Agentforce and Data 360, organizations across healthcare, banking, government, and critical infrastructure face a soberin…OpenAICircles powers telco personalization with OpenAI technologyCircles uses the OpenAI API and Codex to power AI-native telco experiences, increasing ARPU by 22%, reducing churn by 9%, and improving development efficiency.OpenAIHow we built a realtime system for responsive voice AI in six monthsGPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.OpenAIApple is getting this wrongOpenAI addresses Apple’s baseless lawsuit, corrects claims about its employees, and shares messages documenting what happened.NVIDIANVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native StorageStorage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data,...NVIDIAHow to Run Isolated Tenant Kubernetes Clusters on Shared GPU InfrastructureRunning a dedicated Kubernetes cluster per team often results in more isolation than an organization requires. While one cluster can be successfully shared...Microsoft ResearchOrchard: An open framework for scalable agentic AIOrchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to r…MetaGEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation ModelMeta s Generative Ads Recommendation Model (GEM), the foundation model behind ads recommendations across Instagram and Facebook, now trains at LLM scale on several thousand of the latest-generation GPUs. This post goes i…GitLabSecure every commit to production with Claude and GitLabAgentic coding is moving faster than many enterprise governance programs can keep up with. Coding assistants, like the Claude security guidance plugin and Claude Security , can flag and fix common vulnerabilities in code…CloudflareYour agent needs a computer, not a container — introducing @cloudflare/computerAgents need more than just a container to scale. We're introducing @cloudflare/computer, an agent runtime that dynamically orchestrates between fast, efficient isolates and full Linux containers to give every agent a com…CloudflareCloudflare Workers and Containers now support inbound TCP connections and gRPCCloudflare Workers now support inbound TCP connections via Spectrum, allowing direct socket forwarding to Durable Objects and Containers. Developers can run full-duplex gRPC applications or leverage automatic gRPC-to-gRP…CloudflareIntroducing the Billable Usage API: programmatic cost visibility for CloudflareCloudflare has launched a new Billable Usage API for accounts, giving developers and FinOps teams single-endpoint programmatic visibility into cost and usage across all self-serve products. Built around the FOCUS specifi…CloudflareSmaller, faster, safer: running Kimi and GLM at scaleServing frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.CloudflareWorkers RPC now works across Python and JavaScriptOne coding agent can write a Python Worker and another can write a JavaScript Worker. At runtime, those Workers can exchange references to live objects and call their methods without defining APIs, schemas, or serializat…AWSAWS Weekly Roundup: Price reduction of GPT models in Bedrock, CloudWatch managed collectors for Prometheus metrics, and more (August 3, 2026)Last week I had the joy of participating in Amazon’s “Bring Your Kids to Work Day” with my 7 year old son. We commuted together into the New York City office, his first real rush hour train ride, and spent the day explor…AppleUnderstanding Alignment in Multimodal LLMs: A Comprehensive StudyPreference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar to…CloudflareWelcome to Agents WeekAgents Week explores how cloud infrastructure must evolve to serve autonomous agents rather than human browsers. Join us as we unpack the storage, execution, and security primitives needed for an agent-native web.OpenAITen advances in mathematics and theoretical computer scienceOpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.GrabHow AI is transforming analytics at GrabIntroduction At Grab, analytics sits close to almost every decision that matters. Our north star is the democratisation of intelligence, ensuring that anyone making a business call has immediate access to trustworthy ans…CohereLuke RossOpenAIDisrupting a Criminal Scam OperationOpenAI disrupted a Cambodia-based scam operation using ChatGPT to support investment, romance, gambling, and impersonation schemes.OpenAIUnivé builds an AI-ready workforceSee how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovation to transform work at scale.OpenAIBuilding abundant intelligenceA full-stack approach to making advanced AI more capable, more affordable, and more widely useful.OpenAIAdvancing responsible AI across EuropeOpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will continue as the EU AI Act advances.NVIDIANVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate SeekThe demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,...NVIDIACo-Designing AI Model Attention for Fast, Interactive Long-Context InferenceAs agentic and long-context workloads become common, the context lengths increase and attention consumes a larger share of inference time (Figure 1). Because...NetflixModeling Device Capabilities for Analyticsby Aarti Laddha , Richard Diaz-Cool , Rishika Idnani , Venkatesh Selveraj Netflix supports a vast and evolving set of features and content types, ranging from 4K streaming and immersive audio to live streaming and cloud…GitLabHow to govern agentic AI, MCPs, and AI code assistantsAI code completion built human review into the process by design. A developer types, a suggestion appears, and a human decides whether to accept it. A person looked at every line before it shipped. Agentic AI breaks that…GitHubDon’t stop early: Case-folding source code at memory speedHow a branch-free loop and byte-space arithmetic let GitHub case-fold every byte of code search at >45 GiB/s on a single core. The post Don t stop early: Case-folding source code at memory speed appeared first on The Git…EtsyKafka App? There’s a Skill for ThatEtsy is home to over 100 million listings from 5.6 million active sellers. Because the items for sale are unique and creative, there is no standard product catalog that tells us what they are. When someone searches for “…CohereCohere Signs EU AI Content Transparency CodeCloudflareAn API for MoQ: provision your own isolated relaysLast year we made every Cloudflare server a Media over QUIC (MoQ) relay. Now the new provisioning API lets you create your own isolated relay and control who can publish and who can only watch.SalesforceHow Salesforce Built an Agentic Engineering Enablement Strategy for Thousands of Software EngineersGiving thousands of software engineers access to agentic tooling is straightforward. Helping them fundamentally change how they build software is another matter entirely; one every engineering organization eventually run…OpenAIHow avatarin built a 24/7 retail agent with GPT-Realtimeavatarin uses OpenAI’s GPT-Realtime to give Yamada Denki shoppers 24/7 multilingual support. In two weeks, 30,000 people used the agent and 92% of survey responses were positive.OpenAIAdvancing the price-performance frontier with GPT-5.6Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.NVIDIANVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI InfrastructureTwo AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We...NVIDIAFour Ways to Deploy More Secure AI AgentsKnowledge workers are increasingly integrating AI agents into their workflows. Agents that function as "digital coworkers" offer clear benefits. For example,...NVIDIARun High-Performance Core Math at Scale with NVIDIA nvmath-pythonNVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users...NetflixGenRec: Towards LLM-Native Recommendation at NetflixAuthors: Ying Li , Arjun Rao , Shradha Sehgal Introduction Recommendations sit at the heart of the Netflix experience. Our current production models rely on thousands of hand‑crafted features over users, items, and inter…Microsoft ResearchEvoLib: Turning experience into evolving knowledgeLLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib:…Microsoft ResearchEchoverse: Deep, evolving environments for computer-use agentsComputer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the t…Hugging FaceGPU Management: Why Idle GPUs Are the New Grounded AircraftGrabCrowdsourced taxonomy verification: A feedback-driven framework for refining knowledge graph relationships via online search interactionsIntroduction The efficacy of semantic search relies on the accuracy of the underlying Knowledge Graph (KG). In high-velocity domains like on-demand food delivery or e-commerce, the catalog of entities like dishes, produc…GitHubStacked sessions and pull requests in the GitHub Copilot appLearn how I modernized an old codebase of mine using stacked sessions and pull requests in the GitHub Copilot app. The post Stacked sessions and pull requests in the GitHub Copilot app appeared first on The GitHub Blog .Google DeepMindGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaborationGemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.CloudflareDogfooding at scale: migrating cdnjs to Cloudflare’s Developer PlatformWe moved cdnjs, serving 9 billion requests a day, entirely onto Cloudflare's Developer Platform. That means we’re running one of the Internet's busiest open-source CDNs on our own building blocks, and we pushed Workflows…AppleMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action TokenizationTo operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this executi…AppleDimensionality Reduction Meets Network Science: Sensemaking on UMAP’s kNN GraphWhile UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. This…AnthropicInvestigating three real-world incidents in our cybersecurity evaluationsOpenAIHow GPT-5.6 fuses frontier intelligence with frontier efficiencyGPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.OpenAIAccelerating scientific discovery with ChatGPT for Academic ResearchersOpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collaboration, and discovery.OpenAIHow enabling two settings tripled our scores on the ARC-AGI-3 benchmarkHow two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.NVIDIAHow to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo GuardrailsDeploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source...GitLabGitLab Patch Release: 19.2.1, 19.1.3, 19.0.5GitLabWhy GitLab signed the Open Weights and American AI Leadership letterThis week GitLab signed the Open Weights and American AI Leadership letter , joining a long list of other technology companies that support a strong, open AI ecosystem. The letter argues that open weights spur innovation…GitHubTame Dependabot: Group your updates, slow the cadence, keep security fastDependabot keeps your dependencies current, but its defaults can flood your repository with pull requests. Here's how grouping updates, slowing the cadence, and keeping security fixes fast cut the noise on a Microsoft op…DiscordIntroducing Profile Frames: Decorative Borders to Make Your Discord Profile Museum-WorthyIntroducing Profile Frames: a new way to put a finishing touch on your Discord profile. Frames add a decorative border around your profile, giving the whole thing a look that s unmistakably yours.Google DeepMindWe’re launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative controlCloudflarePost-quantum authentication to origins is now supportedCloudflare now supports post-quantum (PQ) authentication when connecting to customer origin servers via Authenticated Origin Pulls and Custom Origin Trust Store. This is the first step towards providing PQ authentication…BAIR (Berkeley)From CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple SiliconFigure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing.…OpenAIScientific computing in the age of agentic AIA new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.NVIDIADeveloping Healthcare Robotics with GPU-Native Medical Physics SimulationUnlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation....Hugging FaceThe OlmoEarth Platform: Geospatial inference at planetary scaleHugging FaceLFM2.5-Encoders for Fast Long-Context Inference on CPUGitHubDisrupting supply chain attacks on npm and GitHub ActionsExplore the changes we've shipped across npm and GitHub Actions over the past few months to disrupt supply chain attack techniques and limit their impact. The post Disrupting supply chain attacks on npm and GitHub Action…Google DeepMindGemini Robotics 2 brings whole body intelligence to robotsCohereWhat Is Agentic AI? Definition and ExamplesCloudflareNatural disasters and government interference: examining Q2 2026’s major Internet disruption eventsCloudflare Radar tracked Internet disruptions driven by natural disasters, government-mandated shutdowns, and DNSSEC key rollovers over the last quarter. This post analyzes traffic telemetry to explain how these events i…AppleMemory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion TransformersSiri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple’s most powerful on-device foundation model. This work presents the memory-efficient a…AirbnbEval-driven development: Lessons from evaluating GenAI at scaleHow Airbnb teams build trustworthy Generative AI products by treating evaluation as a first-class engineering discipline; not an afterthought. Nestled into the lush hillside, this stunning modern retreat features strikin…SpotifyIndexing the Data Lake for Online Point QueriesCompanies like Spotify need vast quantities of data accessible at low latency for online services and,... The post Indexing the Data Lake for Online Point Queries appeared first on Spotify Engineering .SalesforceBuilding Reliable Production AI with Durable WorkflowsBuilding an AI prototype has never been easier. You send a prompt to a model, receive a response, and present the result to a user. As long as each request is independent, the interaction is straightforward to reason abo…PinterestPinner Progression: Better Use-Case Representation Driving Weekly Active User Growth at PinterestPart 1 of 2 Authors Personalization (Homefeed): Yuke Yan, Chuxi Wang, Andreanne Lemay, Olafur Gudmundsson, Anna Kiyantseva, Krystal Benitez, Jongho Kim, Jiacong He, Rahul Goutam, James Li, Dylan Wang User Understanding:…OpenAIHow AI is expanding what people do at workNew OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.NVIDIAAdvancing Semiconductor Innovation Across Materials Engineering and ManufacturingAs AI workloads increase, explosive compute demand is pushing the semiconductor industry to meet unprecedented performance targets. Even small delays can have...NVIDIANVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL CodingModern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware...NVIDIASix Agent Harness Capabilities for Higher Model PerformanceBuilding a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context, executes...NVIDIANVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context LearningNVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they...Hugging FaceAnatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentHugging FaceNVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical RoboticsGitLabClaude Opus 5 on GitLab: Reasoning built for the hard tasksA mistake on a routine task can cost you a minute. A mistake on a large refactor or a debugging trail spanning months of commit history can cost far more, as it compounds silently over hundreds of exchanges. By the time…GitHubThe harness is all you need (mostly)A practical GitHub Copilot workflow for prototyping, planning, implementing, and reviewing software without chasing every new AI tool. The post The harness is all you need (mostly) appeared first on The GitHub Blog .GitHubGitHub Copilot app for Beginners: Getting startedNew to the GitHub Copilot app? Learn how to start projects, work with AI agents, explore canvases, and streamline your development workflow. The post GitHub Copilot app for Beginners: Getting started appeared first on Th…CohereOrchestrate Agentic Workflows with North AutomationsCohereA Day in the Life of a Wealth Manager, With and Without AICloudflareWe’re open-sourcing our privacy proxy CLIpvcli is a curl-like tool designed to simplify the testing of complex privacy protocols like OHTTP.AWSAWS Weekly Roundup: Local Zone in Athens, Claude Opus 5 on AWS, Lambda durable execution for .NET, and more (July 27, 2026)Last week I had the privilege of spending three days in São Paulo with technical builders from across Latin America, brought together for a regional tech event full of deep-dive sessions, hands-on workshops, and conversa…AppleGH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision TasksSystematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Existing slice discovery approaches largely model slices as clusters in repre…AnthropicExpanding our partnership with CognizantAnthropicOur position on open-weights modelsBAIR (Berkeley)Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction.abbel-fig { display: block; text-align: center; margin: 2.4em 0; line-height: 1.4; max-width: 100%; } .abbel-fig img { display: block; margin: 0.65em auto 0; height: auto; max-width: 100%; } /* Image sizes; captions use…NVIDIAModelExpress: Distributing Model Artifacts at the Speed of LightEvery byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly. To make things even worse, moving...GrabAgent platform (Part 1): How we help Grab build and run AI agents at scalePart 1: From one support bot to a framework At Grab, AI agents have evolved from interesting team prototypes into production services used every day by millions of merchants, drivers, and consumers. Today, more than 500…CloudflareBGP ORIGIN attribute manipulation and its impact on the InternetBy doing in-depth testing, we found nearly 70% of BGP paths experience ORIGIN attribute rewrites by transit providers seeking traffic advantages. We examine the global impact of this practice and argue for deprecating OR…AppleLEAD: Breaking the No-Recovery Bottleneck in Long-Horizon ReasoningLong-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate that while decomposition is essential for…AnthropicIntroducing Claude Opus 5SalesforceHow AI Rebuilt Salesforce’s Decades-Old Localization PipelineIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Teresa Marshall, Vice President of Localization. Teresa s team delivers every Salesfo…OpenAILaunching Health in ChatGPTHealth in ChatGPT now lets eligible U.S. users securely connect medical records and Apple Health to get more personalized insights and better understand their health.NVIDIAStart Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in MinutesCustomization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a...NVIDIADebugging Ray Tracing Applications Using NVIDIA OptiX ToolkitNVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways...MicrosoftThe Microsoft 365 Copilot Agent’s Playbook: A Practical Livestream Series for Building Better AgentsBuilding on Microsoft 365 Copilot? Here s your playbook. Declarative agents are quickly becoming one of the most exciting ways to extend Microsoft 365 Copilot and bring organizational knowledge, workflows, and tools dire…Hugging FaceBringing Nunchaku 4-bit Diffusion Inference to DiffusersGitHubThe case for a cooldown: Why Dependabot now waits before issuing version updatesA new default three-day cooldown delays version update pull requests so maintainers and security researchers can address findings in a release before it gets into your code. The post The case for a cooldown: Why Dependab…CloudflareIntroducing Cache Response RulesPerhaps you’ve seen something that should sail out of cache get dragged back to the origin by a stray Set-Cookie or Cache-Control, headers that can be difficult to change on the origin itself. Cache Response Rules is the…PinterestSecuring Infrastructure at Scale: Introducing Pinterest’s Resource Provisioner Pipeline (RPP)Ammar Ekbote | Senior Software Engineer Chan Kim | Senior Software Engineer Managing Infrastructure as Code (IaC) across a massive organization comes with a unique set of security and logistical challenges, particularly…OpenAINTT DATA Group cuts incident analysis to 30 minutes with CodexNTT DATA Group uses ChatGPT Enterprise and Codex to help 9,000 employees automate work, cut incident analysis to 30 minutes, and scale secure AI adoption.OpenAIIntroducing OpenAI PresenceIntroducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.OpenAIAdvancing the next era of national scienceOpenAI outlines its commitment to advancing American science working with the U.S. Department of Energy and national labs to use frontier AI to accelerate discovery.OpenAIHow news organizations are using AI to advance their vital missionsNews organizations are using AI to strengthen reporting, grow audiences, and improve business operations, with OpenAI tools supporting journalists and publishers worldwide.OpenAIBuilding AI infrastructure with the Effingham County communityOpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and access to Codex.NVIDIAMake Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can...GitLabModernize Java with Cursor and GitLabModernize Java 8 to Java 21 sounds like one task. It is not. It touches the build, the runtime, dependencies, APIs, concurrency, tests, containers, and production behavior, often all at once. Ask an agent to do all of th…GitHubNext chapter: Restructuring GitHub’s bug bounty programGitHub is making some significant changes to its bug bounty program, shifting its focus to give researchers a better experience working with the GitHub team. The post Next chapter: Restructuring GitHub s bug bounty progr…GitHubCopilot vs. raw API access: What are you actually paying for?Copilot now bills usage at listed API rates. Compare direct model access with the coding workflow, policy, and harness work around it. The post Copilot vs. raw API access: What are you actually paying for? appeared first…Google DeepMindAccelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis MissionGoogle commits $40M in AI tokens and credits for the Genesis MissionAnthropicA research agenda for the Economic Futures Research FundAnthropicThe Anthropic Economic Index connectorStripeAnalyzing the evidence that helps businesses win “product not received” disputesTo understand what can influence win rates, we analyzed evidence packets from one million disputes over a 16-week period. Here’s what the data shows and what it means for how you mitigate disputes.SalesforceHow AI Reduced Customer Bug Triage from Nearly a Year to Less Than a WeekBy Priya Sethuraman, Abhishek Ghose, Lovish Agarwal, and Aditya Pandey. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priya Sethura…OpenAIDavid Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBCDavid Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC, bringing global leadership in finance, technology, and governance.OpenAIOpenAI and Hugging Face partner to address security incident during model evaluationOpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.OpenAIIntroducing the ChatGPT for small business programOpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.NVIDIANVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AIAgentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context, interact with...NVIDIAInside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AIWhat began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale....NVIDIASetting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...MicrosoftHow to test agent experience changes without shipping themMost changes you think will improve AI agent behavior won't. We tested a dozen hypotheses on a real project upgrade scenario and the majority failed. Learn how to emulate documentation, API, and MCP server changes locall…Hugging FaceGrabette: an open system to record robot-manipulation dataHugging FaceThe State of Simulation for Physical AI: An OverviewGitHubHow to build interactive experiences with canvasesCanvases turn AI into interactive workspaces where you can visualize information, explore workflows, and take action across complex tasks. The post How to build interactive experiences with canvases appeared first on The…Google DeepMindIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberWe’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.Google DeepMindIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash CyberWe’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.CloudflareHow the 2026 World Cup affected Internet trafficWe analyzed global HTTP traffic to explore how kickoff times, streaming habits, and hydration breaks reshaped online activity worldwide. From late-night traffic surges to halftime browsing spikes, here is how the world c…AppleEnvironment-free Synthetic Data Generation for API-Calling AgentsTraining API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments with executable APIs a…AppleAccelerating Text-to-Video Generation with Calibrated Sparse AttentionRecent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. In this paper, we iden…AnthropicDonating another $20 million to Public First ActionAirbnbPersonalizing Airbnb search by learning from the guest journeyHow we built a Transformer-based sequence model that encodes years of guest behavior to surface the right listings at the right time. By: Daochen Zha , Chun How Tan , Xin Liu , Bin Xu , Han Zhao , Xiaowei Liu , Jun Shi ,…SpotifyContent Ingestion & Podcast Video Incident ReportOver the past two months, podcast creators have experienced a series of reliability issues on Spotify. This... The post Content Ingestion Podcast Video Incident Report appeared first on Spotify Engineering .OpenAISafety and alignment in an era of long-horizon modelsOpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.NVIDIAIntegrate NVIDIA Omniverse RTX Sensor Simulation Into Existing AppsDevelopers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and...NVIDIANVIDIA NVLink: The Scale-Up Network for AI FactoriesThe demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute...Hugging FaceIntroducing Cosmos 3 EdgeGitLabGitLab Transcend Hackathon: What developers built on GitLab OrbitWe gave a few thousand developers GitLab Orbit. Then we got out of the way. The community responded with creative solutions to real production problems slowing down their teams. The same problems you hit every week: What…GitLabAutomate work item assignment with a "Work item created" triggerA new, event-driven trigger in GitLab Duo Agent Platform lets flows fire the moment a work item is created, turning triage and assignment from a manual, all-day chore into automation that runs in seconds. This comprehens…GitHub$100 million for open source: A milestone built by the communityCelebrating $100 million contributed by the community to the people who build and sustain open source every day. The post $100 million for open source: A milestone built by the community appeared first on The GitHub Blog…DropboxHow our universal content processing platform Riviera evolved for AI and beyondRiviera is the Dropbox content processing platform that’s been iteratively improving content transformation in our products for roughly a decade.CloudflareCloudflare Internal DNS is now generally availableCloudflare Internal DNS brings authoritative and recursive DNS for private networks to the same global network and control plane that runs Cloudflare's Zero Trust, networking, and public DNS.AWSAWS Weekly Roundup: One-click Lambda setup prompt, OpenAI GPT-5.6 models on Bedrock, and more (July 20, 2026)Last week, my team visited Seoul to meet AWS Korea User Group (AWSKRUG) leaders. AWSKRUG is the largest cloud developer community in Korea, with 20 meetup groups organized by topic and area that collectively host over 10…AppleLVSum: A Benchmark for Timestamp-Aware Long Video SummarizationLong video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantica…AppleLength Value Model: Scalable Value Pretraining for Token-Level Length ModelingToken serves as the fundamental unit of computation in modern autoregressive models, and generation length directly influences both inference cost and reasoning performance. Despite its importance, existing approaches la…AppleRayRoPE: Projective Ray Positional Encoding for Multi-View AttentionWe study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-invariant attention with multi-frequency si…AnthropicApply for Anthropic’s AI for Science rare disease research grantsSalesforceClosing the Loop: How to Build Self-Improving AI Systems with Automated Feedback LoopsAfter six cycles, the system ran out of things to fix. That sentence probably raises more questions than it answers. Here is how we got there. The first skill PR took three days to merge. Twelve comments on structure. Si…OpenAIA scorecard for the AI ageSarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability, and return on compute.NetflixIn-House LLM Serving at NetflixBy AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, in…MicrosoftHow to test agent skills without hitting real APIsYour agent skill calls an API. The moment you start evaluating it, every run either costs money or mutates production data. Learn how to mock APIs transparently so you can run evals without changing your skill or hitting…Hugging FaceFine-tune video and image models at scale with NVIDIA NeMo Automodel and 🤗 DiffusersGitHubThe cost of saying yes has changedThe cost of writing code dropped; the cost of owning it didn't. A framework for deciding which changes are actually cheap in the AI era. The post The cost of saying yes has changed appeared first on The GitHub Blog .Google DeepMindIntroducing Gemini 3.5 Flash CyberGoogle introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.CloudflareCloudflare WAF protects WordPress applications from two high-severity vulnerabilitiesCloudflare has deployed two WAF rules in response to high-severity vulnerabilities disclosed to us by the WordPress security team. The new rules protect all Cloudflare customers using affected WordPress versions, but cus…AppleShow Me Examples: Inferring Visual Concepts from Image SetsVision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, current models fail to infer shared concepts from sets of example images and a…AppleWhen Unlearning Is Free: Leveraging Low Influence Points to Reduce Computational CostsAs concerns around data privacy in machine learning grow, the ability to unlearn—or remove—specific data points from trained models becomes increasingly important. While state-of-the-art unlearning methods have emerged i…YelpMigrating from Apollo Tooling to GraphQL Codegen at YelpIntroduction At Yelp, we rely heavily on GraphQL and Apollo for data loading in our frontend React monorepo. When a developer writes a GraphQL query or mutation inside a React component, the shape of the response is defi…SalesforceUsing Claude to Build an AI Knowledge Base in 30 MinutesEver had an experience onboarding someone, where you had to provide them a handful of documents, all authored at different times, and in various states of completeness? Ever had your AI agent burn cycles pulling down tha…OpenAIHow Cars24 scales conversations and builds faster with OpenAICars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company.OpenAIHow Codex became a collaborator for OpenAI’s creative teamHow OpenAI’s creative team uses Codex to build custom creative tools, accelerate ideation, and prototype faster with context-aware AI.OpenAIWhy teens deserve access to safe AILearn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.NVIDIAScaling Agentic AI Factories Through Extreme Co-Design with NVIDIA BlueFieldAgentic AI changes the infrastructure pattern for AI factories. One request can trigger many model calls, tool calls, memory lookups, policy checks, storage...NVIDIAIntegrating Context-Aware Video AI Agents Into Enterprise WorkflowsA video analytics AI agent that can perceive, reason, and act based on massive amounts of video footage must be integrated with existing workflows and...NVIDIAQ&A: How Capcom Brought Path Tracing to RE ENGINE Across PRAGMATA and Resident Evil RequiemCapcom's RE ENGINE team set out to bring path tracing into two shipping titles at once, Resident Evil Requiem and PRAGMATA, each with a different visual...Hugging FaceSecurity incident disclosure — July 2026Hugging FaceNewer Models, Same AdvantageHugging FaceNVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic RetrievalGitLabWhen a version bump breaks your build, GitLab fixes itAI is writing more code and pulling in more dependencies, increasing application risk. Most of that exposure isn t from code your team actively chose. A 2025 study of the Maven ecosystem found vulnerabilities reaching ro…GitLabGitLab 19.2 releasedGitLabForrester Consulting: GitLab Duo Agent Platform delivers 400% ROIA new Forrester Consulting Total Economic Impact™ study found that organizations using GitLab Duo Agent Platform achieve a 400% return on investment and $7.5 million in net present value over three years — with payback i…GitLabBring GitLab Duo Agent Platform to your terminalMost of the work for software delivery doesn’t happen only in the editor. Pipelines fail. Tests break. Vulnerabilities show up. And a lot of that work starts and ends at the command line. Agentic AI in the terminal that…GitLabGitLab Duo Security Review spots logic flaws scanners missStatic scanners excel at catching vulnerabilities that fit a known pattern, like unsanitized query inputs, hardcoded secrets, and unsafe deserialization. They struggle against flaws in your application’s logic, where the…GitLabTurn multi-step software delivery into agentic flows you can trustKnowing what to do next in software development is rarely the hard part. Doing it again in the exact same steps — implement an issue, fix a pipeline, review a merge request — is. Chat that only provides answers still lea…Google DeepMindOur approach to bioresilienceGoogle DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.CohereCohere and University of Toronto Partner on Responsible AIAppleLocation-Invariant Properties of Functions Versus Properties of Distributions: United in Testing but Separated in VerificationA property of functions is called location-invariant (or symmetric) if it can be characterized in terms of the frequencies in which each value occurs in the function, regardless of the locations in which each value occur…AppleDoubly Sub-linear Interactive Proofs of ProximityWe study doubly sub-linear interactive proofs of proximity (dsIPPs): proofs that are ultra-fast to generate, and can be used to prove approximate assertions about a huge input. Proof generation is ultra-fast in the sense…ApplePersonalizing Incremental Video Search with Hybrid Text and ID EmbeddingsIncremental video search requires high-quality ranking after each keystroke, where intent is often underspecified (e.g., 1–3 character prefixes). We present a personalization system for Apple TV search that combines comp…AppleEmbarrassingly Simple Self-Distillation Improves Code GenerationCan a large language model (LLM) improve at code generation using only its own raw outputs, without a verifier, a teacher model, or reinforcement learning? We answer in the affirmative with simple self-distillation (SSD)…AppleInteractive Proofs for General Distribution PropertiesSuppose Alice has collected a small number of samples from an unknown distribution, and would like to learn about the distribution. Bob, an untrusted data analyst, claims to have run a sophisticated data analysis on the…OpenAIGPT-Red: Unlocking Self-Improvement for RobustnessExplore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.OpenAIThe US is advancing AI safety through state and federal actionOpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.NVIDIABuilding Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3For over fifteen years, x86 CPUs have shipped with a dedicated hardware instruction for carryless multiplication. It’s a small but stubborn primitive that...NVIDIADevelop Lightweight USD Runtimes Faster with AI AgentsOpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation...NVIDIABuild a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 SkillsDevelopers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking...Mistral AIRobostral Navigate: single-camera AI navigationMicrosoftBuilding AX evals that actually workThis is the eighth and final article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent sta…MicrosoftBuilding AX evals that actually workThis is the eighth and final article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent sta…MetaExploring Hierarchical Interest Representation For Meta Ads Deep Funnel OptimizationHierarchical Interest Representation is a research area for Meta Ads. We’re exploring an upstream representation layer over the universe of Ads entities – users, advertisers, products, services – learning unified embeddi…Hugging FaceIntroducing Real World VoiceEQ: Measuring the human quality of voice AIHugging FaceWelcome Inkling by Thinking MachinesHugging FaceModel Routing Is Simple. Until It Isn’t.Hugging FaceWhat building Shippy taught us about building agentsGitHubGitHub for Beginners: Your roadmap to mastering the GitHub essentialsNew to GitHub? This beginner's guide explains version control, repositories, and pull requests—plus everything else you need to start working confidently on GitHub. The post GitHub for Beginners: Your roadmap to masterin…CohereThe Total Cost of AI Ownership (AI TCO)CohereLanguage, Decoded: What if you could code in your own language?CohereAriana MilliganAppleUncertainty Quantification for LLM Function-CallingLarge Language Models (LLMs) are increasingly deployed to autonomously solve real-world tasks. A key ingredient for this is the LLM Function-Calling paradigm, a widely used approach for equipping LLMs with tool-use capab…AppleCLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningRetrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge but still suffers from long contexts and disjoint retrieval–generation optimization. In this work, we propose CLaRa (Cont…AppleOne Layer Is Enough: Adapting Pretrained Visual Encoders for Image GenerationVisual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in leveraging high-quality p…YelpTraining Orchestrator: Unifying Model Training at YelpAt Yelp, we train many machine learning models on different schedules. Applied machine learning teams all have their own set of Spark-based training batches, scripts, and configurations. Over time, these diverged, leadin…SlackShipyard: How We Built Slack’s Next-Generation EC2 PlatformOver the past few years, we’ve been on a journey to modernise how we run Amazon Elastic Compute Cloud (EC2) instances at Slack. In our first post, Advancing Our Chef Infrastructure, we shared how we moved from a single C…OpenAIHow sales teams use ChatGPT WorkSee how sales teams can use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, account plans, and stalled-deal diagnoses from real work inputs.OpenAIHow data science teams use ChatGPT WorkSee how data science teams can use ChatGPT Work to build root-cause briefs, impact readouts, KPI memos, scoped analyses, and dashboard specs from real work inputs.OpenAIHow to manage AI investments in the agentic eraLearn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.NVIDIAPost-Train NVIDIA Cosmos 3 in One Day Using Agent SkillsWhat if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning...NVIDIAHow to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMoCoding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve...NVIDIALessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI ReasoningThe NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when...InstacartBlueberry: Force Multiplier For The On-Call EngineerHow we built a Slack-native on-call reasoning harness at Instacart that shortens time to first insight, speeds up theory testing, and turns tribal knowledge into reusable infrastructure. Key Contributors: Karthik Halukur…CohereTiny Aya Expedition Drives Multilingual InnovationCohereFrom Reviewing Words to Building Africa's AI: Our Journey Inside the Cohere Labs Open Science CommunityCloudflareA broken DNSSEC rollover took down .al. Now 1.1.1.1 tells you when validation is bypassedWhen a failed DNSSEC key rollover took down the .al TLD, we deployed a Negative Trust Anchor to restore resolution. This time, though, clients didn't have to take our word for it: 1.1.1.1 returned EDE 33, a new DNS error…AppleProactive Agent Research Environment: Simulating Active Users to Evaluate Proactive AssistantsProactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Existing approaches mod…AppleMultilingual Semantic Retrieval for Apple Music SearchApple Music serves listeners across 150+ storefronts in dozens of languages, with a catalog that grows by hundreds of thousands of new tracks daily. At this scale, search recall on misspelled, transliterated, and cross-l…AnthropicAnthropic commits $10 million to Canadian AI researchAnthropicIntroducing Claude for TeachersAirbnbFrom weeks to a day: how we made LLM evaluation fast enough to iterate onTraining an LLM is the easy part. The hard part is designing experiments and evaluations that you can trust enough to know whether the new model is actually an improvement. By : Baharak Saberidokht Introduction Shipping…NVIDIAExtreme Event Likelihoods with Guided Generative ModelsAcross science, engineering, and finance, many of the most important risks come from low-likelihood, high-impact events. Estimating the probability of these...NVIDIANVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300xUseful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes...NetflixBuilding Service Topology at Scale: Architecture, Challenges, and Lessons LearnedBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez-Silva Nathan Fisher A deep dive into the engineering challenges of building a real-time service dependency map at Netflix scale: from streaming architectures a…Microsoft ResearchVerifying Rust cryptography in SymCrypt, from standards to codeCryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The p…MetaModernizing the Meta Ads Service With an Open-Source Kernel SchedulerTL; DR At Meta s scale, a few milliseconds of latency degradation can have a significant negative impact on ads performance. When a Linux kernel upgrade risked regressing latency across Meta s ad serving fleet, we turned…Google DeepMindEmpowering India’s next generation of innovators with ATL SaathiGoogle and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.CloudflareIntroducing Precursor: detecting agentic behavior with continuous client-side signalsPrecursor, our new continuous behavioral validation engine for bot management, offers visibility into how humans and bots actually interact across the full user journey. By turning session-level behavior into bot detecti…AWSAWS Weekly Roundup: AWS Builder Center at 1 year, Network Scanning in Security Hub, Loom for AWS, and more (July 13, 2026)AWS Builder Center turned one year old last week. Launched on July 9, 2025, the platform has grown from a community hub with Wishlist voting, community profiles, and a toolbox into a full ecosystem with sandbox environme…AWSAmazon SQS turns 20: Two decades of reliable messaging at scaleOn July 13, 2006, we launched Amazon Simple Queue Service (Amazon SQS) as one of the first three services available to customers, alongside Amazon EC2 and Amazon S3. We had learned firsthand that distributed systems need…NVIDIAHow to Evaluate General-Purpose Robot Policies for Real-World DeploymentRobotics foundation models have made remarkable progress. Today's best systems can follow natural language instructions to pick, place, sort, and manipulate a...OpenAIHow Deutsche Telekom is rewiring telecommunications with AIHow Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and the future of voice.NVIDIAAccelerating End-to-End Co-Folding Performance with NVIDIA BioNeMo Agent ToolkitBiomolecular structure prediction and co-folding with models like OpenFold3 are now mainstream, large-scale workloads powering drug discovery and protein...NVIDIAAI Model Co-Design: Hardware-Friendly LLM DesignAI performance comes down to three dimensions: Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a...NVIDIAKernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch OverheadThere are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead,...NVIDIAReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host OffloadingLarge language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,...Hugging FaceProfiling in PyTorch (Part 3): Attention is all you profileGrabScaling Grab's Data Lake: Our journey to Apache Iceberg adoptionIntroduction: The evolution of Grab’s Data Lake At Grab’s scale, managing petabytes of data across billions of S3 objects demands more than a storage layer. It demands a robust architectural primitive that supports the h…GitHubBetter tools made Copilot code review worse. Here’s how we actually improved it.How migrating Copilot code review to shared Unix-style code exploration tools reduced review cost by reshaping agent workflows around pull request evidence. The post Better tools made Copilot code review worse. Here s ho…CohereHardware-Aware, Dynamic Speculative Decoding (DSD)CohereMulti-Agent Systems: Enterprise AI GuideCloudflareImproving Smart Tiered Cache for Public Cloud RegionsSmart Tiered Cache allows for precise upper tier selection for origins hosted on AWS, GCP, Azure, and Oracle Cloud with customer-provided cloud region hints.AppleBehavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized PoliciesThis paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and Security (ARES) 2026. Autonomous…SalesforceHow Informatica Reduced Data Integration Pipeline Development from Days to MinutesIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Nancy Chen, Vice President of Engineering at Informatica. Nancy leads the development…NVIDIAA Practical Guide to GPU-Initiated Communication for Molecular Dynamics at ScaleMolecular dynamics (MD) simulations are among the most demanding workloads in computational science. Using them, researchers can observe atomic behavior in...NVIDIASynthetic Data Generation for Financial AI Research with NVIDIA NeMoFine-tuning LLMs for financial natural language processing (NLP) is constrained by limited, imbalanced data. Real-world financial news overrepresents earnings...Microsoft ResearchAurora 1.5: Extending open foundation models for weather and Earth-system applicationsAurora 1.5 adds 22 more variables, hourly temporal resolution, and probabilistic ensemble forecasting to the Aurora foundation model, making it more useful for real-world weather, climate, and energy applications. The po…Mistral AIVersion control for prompts & skills in StudioLyftFrom Day 1 to Production: Building Lyft’s Analytics & Rides Intelligence Assistant as Onboarding…From Day 1 to Production: Building Lyft’s Analytics Rides Intelligence Assistant as Onboarding Project Written by Sagar Baronia at Lyft. A Different Kind of Day One Most onboarding journeys follow a familiar arc: orienta…GitLabGreen DevOps: Why carbon measurement belongs in your CI/CD pipelineA typical software team runs hundreds of CI/CD jobs a day. Each one runs on compute and burns energy that doesn t show up in your pipeline logs, including its carbon impact. That invisibility is exactly the problem. You…GitHubHow GitHub gave every repository a durable ownerGitHub had over 14,000 repositories. Fewer than half had clear ownership. Here's how we gave every active repository a validated owner in under 45 days, archived the rest, and made ownership the foundation for everything…CloudflareWhy we cannot wait for better post-quantum signature algorithmsNIST is advancing nine new post-quantum signature algorithms as potential candidates for future standardization. We take a closer look at all of them, and argue that while they are in the works and show great potential,…AppleUnmasking On-Policy Distillation: Where It Helps, Where It Hurts, and WhyOn-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental. Which teacher mode…AppleRecursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long ContextLong-context handling remains a core challenge for language models: even with extended context windows, models often fail to reliably extract, reason over, and use the information across long contexts. Recent works like…AppleIncentivizing Temporal-Awareness in Egocentric Video Understanding ModelsMultimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct…AnthropicIntroducing a way to reflect on how you use ClaudeAnthropicInviting hard questionsAnthropicBen Bernanke appointed to Anthropic’s Long-Term Benefit TrustAnthropicUST is bringing Claude to physical AINVIDIARunning Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance...NVIDIACreate a LangChain Deep Agents Harness Profile for NVIDIA Nemotron 3 Ultra to Improve PerformanceAgentic systems often face a trade-off between accuracy and cost. The highest-performing proprietary frontier models and harnesses provide top accuracy but are...Microsoft ResearchFlint: A visualization language for the AI eraShort chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, huma…MicrosoftThe hidden variables in your agent evalThis is the seventh article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how t…MicrosoftLet’s Learn GitHub Copilot App – Free Virtual Training EventJoin us for a free online event series kicking off July 16 to learn how to get started with the GitHub Copilot App! The post Let’s Learn GitHub Copilot App Free Virtual Training Event appeared first on Microsoft for Deve…MicrosoftThe hidden variables in your agent evalThis is the seventh article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how t…MicrosoftLet’s Learn GitHub Copilot App – Free Virtual Training EventJoin us for a free online event series kicking off July 16 to learn how to get started with the GitHub Copilot App! The post Let’s Learn GitHub Copilot App Free Virtual Training Event appeared first on Microsoft for Deve…Hugging FaceNative-speed vLLM transformers modeling backendHugging FaceData for AgentsGitLabHow we used AI agents to migrate GitLab rate limitingA small team at GitLab spent the past few weeks running an experiment: Could we use AI agents to migrate part of our legacy rate-limiting system without dropping the safety bar? Short answer: yes. AI agents do work. They…GitLabGitLab Patch Release: 19.1.2, 19.0.4, 18.11.7GitHubHow GitHub Copilot enables zero DNS configuration for GitHub PagesGo from an empty repository to a live custom domain with HTTPS in about 14 minutes, without manually editing a single DNS record. The post How GitHub Copilot enables zero DNS configuration for GitHub Pages appeared first…GitHubGitHub availability report: June 2026In June, we experienced six incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: June 2026 appeared first on The GitHub Blog .GitHubAutomating cross-repo documentation with GitHub Agentic WorkflowsExplore how the Aspire team turns merged product changes into SME-reviewed docs pull requests, closing the gap between release and documentation. The post Automating cross-repo documentation with GitHub Agentic Workflows…CloudflareIntroducing Meerkat: an experiment in global consensusCloudflare Research is building a global consensus service called Meerkat that uses a new consensus algorithm called QuePaxa. We plan to use Meerkat to build a strongly consistent, fault-tolerant key-value store, and oth…NVIDIABuilding an Analysis AI Agent for Industrial Alarm Management with NVIDIA NemotronIndustrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context,...NVIDIAMaximize Spectral Efficiency with AI-Native RAN and NVIDIA AI AerialSpectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to...NVIDIADevelop Humanoid Robot Policies End-to-End with NVIDIA Isaac GR00TAs more teams move from humanoid robot bring-up to task-specific skill development, the need for repeatable development workflows is growing. Building humanoids...NVIDIANVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic WorkloadsAgentic systems turn model reasoning into action through multi-step workflows that combine inference, tool use, code execution, retrieval, orchestration, and...MicrosoftDon’t rewrite your CLI for agentsThere s advice making the rounds: replace your CLI args with a single --json payload so agents can use your tool more effectively. The thinking being, that agents already think in structured formats, and nested data maps…MicrosoftDon’t rewrite your CLI for agentsThere s advice making the rounds: replace your CLI args with a single --json payload so agents can use your tool more effectively. The thinking being, that agents already think in structured formats, and nested data maps…Hugging FaceLeRobot v0.6.0: Imagine, Evaluate, ImproveHugging FaceRun AI workloads on any cloud, store on Hugging Face: zero-egress storage with SkyPilotHugging FaceHugging Face Models on Foundry Managed ComputeHugging FaceFrom Hugging Face to Amazon SageMaker Studio in one clickGitHubQ1 2026 Innovation Graph update: Open source collaboration is accelerating worldwideNew Innovation Graph data shows global developer communities growing faster than ever, with collaboration reaching new highs across many economies. The post Q1 2026 Innovation Graph update: Open source collaboration is a…DiscordDiscord Patch Notes: July 7, 2026Check out the finer details of the more technical fixes implemented into Discord recently.CohereCohere Transcribe Arabic: Open-Source Speech AICohereInside Tiny Aya: What an SAE Sees in a Model Built for 70+ LanguagesCloudflareCloudflare proudly joins the UK government's Cyber Resilience PledgeThe pledge is a voluntary framework inviting organizations to commit to foundational cyber security governance, board-level accountability, and supply chain rigor. For over a decade, Cloudflare has pioneered the core pil…BAIR (Berkeley)Intelligence is Free, Now What? <br> Data Systems for, of, and by Agents... government of the people, by the people, for the people ... Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost roughly $30 per million tokens in early 2023; t…AppleTaming Text-to-Sounding Video Generation via Advanced Modality Condition and InteractionThis study focuses on Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text, with both modalities aligned to the text conditions. Despite progress in joint audio-video…AppleFlowEval: Reference-Based Evaluation of Generated User InterfacesWhile large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction design. Existing eval…AppleWeblica: Scalable and Reproducible Training Environments for Visual Web AgentsThe web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories for supervised fine-tu…AppleA Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language ModelsSafety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expressed, and concept neurons that encode the harmful knowledge itself. B…AppleDynaMiCS: Fine-Tuning LLMs with Performance Constraints Using Dynamic MixturesMulti-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, instruction following, or safety evaluati…AppleLensVLM: Selective Context Expansion for Compressed Visual Representation of TextVision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to…AppleMT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow MatchingRecent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, edi…SalesforceBuilding Enterprise AI Agents That Are Both Autonomous and ReliableWhile red-teaming an early refund agent early last year, one of our engineers typed a deliberately absurd line: my very valid and very verified email. The agent, running on a frontier LLM, accepted it as proof of identit…NVIDIAEnhancing Goodput in Large-Scale LLM Training with Nonuniform Tensor ParallelismTraining LLMs at massive scale brings unique infrastructure challenges, especially as jobs span thousands of GPUs and run for extended periods. The longer these...MicrosoftNot all model upgrades are upgradesA new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks a…MicrosoftNot all model upgrades are upgradesA new model drops with lower per-token pricing and better benchmarks. You switch. A week later someone asks why the agent is burning 12x more tokens on the same task while producing worse output. We ran 150 agent tasks a…Hugging Face🤗 Kernels: Major UpdatesHugging FacePRX Part 4: Our Data StrategyGitLabKeep your GitLab seats in check with restricted accessGitLab restricted access for instance admins, group owners, and billing managers enables predictable seat costs with less manual gatekeeping. The feature has been significantly improved and is now more complete for the w…CloudflareYour Worker can now have its own cache in front of itWe are launching Workers Cache, a regionally tiered cache that sits directly in front of your Worker entrypoints. Infinitely composable, configured via standard HTTP headersAWSAWS Weekly Roundup: Claude Sonnet 5 on AWS, Amazon WorkSpaces for AI agents, AWS service availability updates, and more (July 6, 2026)A couple of editions ago I wrote about what I find so energizing about working with startups. Last week I got a fresh dose of it: I spent a few days with the AWS Startups team, listening to stories of founders talking ab…AppleScaling Properties of Continuous Diffusion Spoken Language ModelsSpeech-only spoken language models (SLMs) lag behind text and text-speech models in performance, with recent discrete autoregressive (AR) SLMs indicating significant computational and data demands to match text models. S…AppleUnderstanding Annotator Safety Policy with InterpretabilitySafety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such as operational fail…AppleSegmental Attention Decoding with Long Form Acoustic EncodingsWe address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to encode absolute frame positions by exploit…AppleFortress: A Case Study in Stabilizing Search Recommendations via Temporal Data Augmentation and Feature PruningIn search and recommendation systems, predictive models often suffer from temporal instability when certain input features introduce volatility in output scores. This instability can degrade model reliability and user ex…AppleRevisiting ASR Error Correction with Specialized ModelsLanguage models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models (LLMs) have been applied to ASR correcti…ApplePath-Constrained Mixture-of-ExpertsSparse Mixture-of-Experts (MoE) architectures route each token through a subset of experts at each layer independently. We propose viewing MoE computation through the lens of expert paths—the sequence of expert selection…AppleTopoPrimer: The Missing Topological Context in Forecasting ModelsWe introduce TopoPrimer, a framework that makes the global topological structure of the series population an explicit input to any forecasting model. TopoPrimer improves accuracy across diverse domains, stabilizes foreca…AnthropicGovernment of Alberta uses Claude to find and fix cybersecurity vulnerabilities across government systemsGrabMigrating Counter Service storage: Design choices and learningsIntroduction Counter Service is used across Grab’s anti-fraud platform to answer time-windowed count questions, such as recent ride requests by a user or failed payment attempts on a card. The service handles tens of tho…Google DeepMindGoogle DeepMind and A24 announce first-of-its-kind research partnershipCohereSinhala Is Not Just Low-Resource: It Is Under-EvaluatedMistral AILeanstral 1.5: Proof Abundance for AllGitHubHow GitHub used secret scanning to reach inbox zeroGitHub had 20,000+ secret scanning alerts across 15,000 repositories. Here's how we separated signal from noise, built remediation workflows, and reached inbox zero in nine months. The post How GitHub used secret scannin…AppleResidual Context Diffusion Language ModelsDiffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel. However, state-of-the-art block-wise dLLMs rel…AppleLearning Unmasking Policies for Diffusion Language ModelsDiffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference. One critical design a…AppleLearning Structured Reasoning via Tractable Trajectory ControlLarge language models can exhibit emergent reasoning behaviors, often manifested as recurring lexical patterns (e.g., “wait,” indicating verification). However, complex reasoning trajectories remain sparse in unconstrain…AnthropicMore details on Fable 5’s cyber safeguards and our jailbreak frameworkSalesforceHow AI Learned to Investigate Mobile Build Failures Like an Experienced Support EngineerIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Archana Indran, Senior Engineering Manager for Mobile CI/CD. Archana s team built Ana…MicrosoftWhat AI benchmarks are not telling youThis is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MicrosoftWhat AI benchmarks are not telling youThis is the sixth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MetaMeta’s AI Storage Blueprint at ScaleOver the past several years, model capabilities and training dataset sizes have experienced exponential growth. During the past year or so, the time between new-frontier-model releases has gone down from months to weeks.…InstacartVariance Reduction Below the Randomization GrainSergio Camelo, Caitlin Kearns, Matias Cersosimo, and Tilman Drerup As artificial intelligence increases the velocity of engineering and science teams, experimental throughput is set to become a bottleneck for many produc…Hugging FaceHugging Face and Cerebras bring Gemma 4 to real-time voice AIGitLabGitLab Patch Release: 18.8.11GitHub6 security settings every GitHub maintainer should enable this weekThese six free settings will not make your project unhackable. Nothing will. What they will do is close the easy doors. Turn these on, and your project will be meaningfully harder to attack than it was before. The post 6…CloudflareContent Independence Day, one year on: building the business model for the agentic InternetOne year after declaring Content Independence Day, a dynamic market for monetized content has officially emerged. In this report, we examine how the rise of autonomous AI agents is upending traditional search referrals a…CloudflareAnnouncing the Monetization Gateway: charge for any resource behind Cloudflare via x402We're opening the waitlist for our Monetization Gateway, which will allow you to charge for any web page, dataset, API, or MCP tool behind Cloudflare. The charges will settle in stablecoins over the x402 open protocol, w…BAIR (Berkeley)2026 BAIR Graduate ShowcaseCongratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed…AWSUpgrade Amazon EKS clusters with confidence using Kubernetes version rollbacksLearn how Kubernetes version rollbacks for Amazon EKS let you reverse cluster upgrades within seven days. This new feature provides a safety net for upgrade failures—no cluster rebuilds required—turning Kubernetes versio…Microsoft ResearchSkillOpt: Agent skills as trainable parametersAI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process, making agent behavior more reliable with…Meta10 Years of Meta’s Commitment to PythonThis year marks Meta s 10th consecutive year as a sponsor of the Python Software Foundation (PSF), the charitable organization dedicated to advancing, supporting, and protecting the open-source Python programming languag…Hugging FaceFeaturing Every Eval Ever Results on Hugging Face Model PagesHugging FaceWhy Specialization Is InevitableHugging FaceScarfBench: Benchmarking AI Agents for Enterprise Java Framework MigrationGitLabClaude Sonnet 5 on GitLab: More reliable, more efficientAnthropic’s Claude Sonnet 5 is now available on GitLab Duo Agent Platform across all tiers and deployment models through GitLab s AI Gateway. Claude Sonnet 5 is built for work that agents assist software teams with every…GitHubHow GitHub maintains compliance for open source dependenciesExplore how the Open Source Program Office uses GitHub’s new license compliance product to manage open source dependencies at scale. The post How GitHub maintains compliance for open source dependencies appeared first on…DiscordCost Attribution in Discord’s APIDiscord s API spans 1700+ endpoints across hundreds of Kubernetes deployments. The challenge: tracking per-feature hosting costs without restructuring. Jim Benton helps explain how Discord tackled the situation.DiscordDiscord is Now on Meta Quest: Reach Out to Your Servers While in VRStarting today, Discord is now available to download directly from Meta’s Horizon App Store. No more sideloading. No more using the web app. Grab it using the link in this blog, or at discord.com/download.Google DeepMindStart building with Nano Banana 2 Lite and Gemini Omni FlashAWSAutomate public TLS certificate issuance with ACME support in AWS Certificate ManagerAWS Certificate Manager now supports the ACME protocol for public TLS certificates, enabling automated issuance and renewal through any ACMEv2-compatible client on any workload. Administrators get centralized governance,…AWSAmazon EC2 C9g and C9gd instances powered by AWS Graviton5 processors are now availableAmazon EC2 C9g and C9gd instances, powered by AWS Graviton5, are now generally available. They deliver up to 25% better compute performance than Graviton4-based instances, 5x larger cache, fastest memory of any processor…AWSAccelerate your infrastructure deployments by up to 4x with AWS CloudFormation Express modeAWS CloudFormation speeds up infrastructure deployment with Express mode, enabling AI agents and developers to receive deployment confirmation in seconds and iterate faster. Available in all commercial Regions at no addi…AnthropicRedeploying Claude Fable 5AnthropicIntroducing Claude Sonnet 5AnthropicClaude Science, an AI workbench for scientists, is now availableSalesforceInside Unified Planner: The AI Brain Behind AgentforceIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Gaurav Aggarwal, a Software Engineering Architect. Gaurav developed the Unified Plann…NetflixGenPage: Towards End-to-End Generative Homepage Construction at NetflixAuthors: Lequn Wang , J iangwei Pan , and Linas Baltrunas Figure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a time, each one conditioned on what’s already on the page an…Microsoft ResearchMemora: A Harmonic Memory Representation Balancing Abstraction and SpecificityAI agents can't remember past conversations. They must constantly reload or retrieve context, which grows less efficient as tasks get longer and more complex. Memora solves this with a scalable memory system separating w…InstacartLeveraging PyFixest for High-Cardinality Marketplace Modeling at InstacartBenjamin S. Knight Scaling Marketplace experiments requires specialized statistical techniques. We examine why standard ordinary least squares regression (OLS) becomes computationally intractable when controlling for hig…Hugging FaceDiScoFormer: One transformer for density and score, across distributionsGitLabWhat's new in Git 2.55.0?The Git project recently released Git 2.55.0 . Let s look at a few notable highlights from this release, which includes contributions from the Git team at GitLab. What s covered: git-history(1) learns fixup fsmonitor dae…GitHubInside the Advisory Database and what happens when vulnerability volume breaks recordsThe GitHub Advisory Database is processing more vulnerability reports than ever before. Here's what's driving the surge, how we're responding, and how the community can help. The post Inside the Advisory Database and wha…GitHubHighlights from Git 2.55The open source Git project just released Git 2.55. Here is GitHub’s look at some of the most interesting features and changes introduced since last time. The post Highlights from Git 2.55 appeared first on The GitHub Bl…AWSAWS Weekly Roundup: Agentic CX designer for Amazon Connect Customer, EC2 AMI Watermarks, Open Governance for MySQL, and more (June 29, 2026)It has been a busy stretch on the AWS Summit circuit. At the New York City Summit, I delivered a workshop called Building AI architectures with AWS Serverless, and it was a lot of fun watching builders wire up agents and…MicrosoftYour agent already has a planIf an agent isn t doing the right thing, the obvious move is to make the docs clearer. Add a tip, spell out the correct command, describe the right approach more prominently. You do all of that, and the agent still ignor…MicrosoftYour agent already has a planIf an agent isn t doing the right thing, the obvious move is to make the docs clearer. Add a tip, spell out the correct command, describe the right approach more prominently. You do all of that, and the agent still ignor…Hugging FaceRun a vLLM Server on HF Jobs in One CommandGitHubTransitioning as a HubberHow GitHub's culture and benefits helped me be the best version of myself. The post Transitioning as a Hubber appeared first on The GitHub Blog .GitHubGitHub and UNDP team up to advance development priorities in Ghana with open sourceGitHub joined the United Nations Development Programme in Ghana to explore how open source governance can support one of West Africa's most ambitious digital reform efforts. The post GitHub and UNDP team up to advance de…PinterestAchieving Near-Linear Training Scalability for Pinterest’s Foundation ModelsSheng Huang | Software Engineer, AI Platform; Pong Eksombatchai | Machine Learning Engineer, Applied Sciences; Saurabh Vishwas Joshi | Software Engineer, AI Platform; Gaurav Arora | Software Engineer, AI Platform; Karthi…Microsoft ResearchUnderstanding the brain with AI-driven explanations and experimentsResearchers introduce generative causal testing, which translates black box models into clear hypotheses and verifies them in the scanner, revealing what specific brain regions respond to in language. The post Understand…MicrosoftLearn from Microsoft: Transform software development through an agentic platformSee how Microsoft is transforming software development with agentic workflows, AI-powered automation, and specialized agents across the engineering lifecycle. The post Learn from Microsoft: Transform software development…MicrosoftLearn from Microsoft: Transform software development through an agentic platformSee how Microsoft is transforming software development with agentic workflows, AI-powered automation, and specialized agents across the engineering lifecycle. The post Learn from Microsoft: Transform software development…MetaPrivacy-Aware Infrastructure in the AI-Native Era: An Asset Classification Case StudyPrivacy controls — systems that enforce retention, access, allowed-purpose, downstream-sharing, or anonymization policies — require a reliable understanding of data to function. Before such a control can operate effectiv…Hugging FaceWhich tokens does a hybrid model predict better?GitLabGoogle Antigravity agents get full context with GitLab OrbitDevelopers working in Google Antigravity can now install our lifecycle context graph, GitLab Orbit , directly from the Antigravity MCP Store and give their agents structured access to projects, pipelines, merge requests,…GitHubEvaluating performance and efficiency of the GitHub Copilot agentic harness across models and tasksExplore how the GitHub Copilot agentic harness delivers strong results across multiple benchmarks and leading token efficiency, while maintaining flexibility to choose among more than 20 models. The post Evaluating perfo…DropboxHow we used DSPy to turn AI evaluations into better responses in Dash chatWe used DSPy to improve LLM judges and optimize our chat experience, creating an evaluation-driven feedback loop that produced better outputs.DiscordDiscord Update: June 25, 2026 ChangelogHere s the Discord Changelog from June 25, 2026, so you can stay informed on what’s new in recent app updates!CohereAutomating Fork Maintenance with AI AgentsCohereThe Solution: North becomes a security agentPinterestAutomated Schema Evolution in Pinterest’s Next-Generation DB Ingestion FrameworkYisheng Zhou | Software Engineer II Liang Mou | Sr Staff Software Engineer Gabriel Raphael Garcia Montoya | Staff Software Engineer Istvan Podor | Staff Software Engineer Introduction In the first post of this series , w…Microsoft ResearchTalos: Scaling rare disease diagnosis with automated, iterative genomic reanalysisTalos was built to help resolve a major bottleneck in genomic medicine: human review time. The open-source system recovered 90% of in-scope diagnoses while surfacing just 1.3 candidate variants per patient for expert rev…Mistral AIBringing more control over your connectorsMicrosoftWhen the model has never seen your codeThis is the fifth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…MicrosoftWhen the model has never seen your codeThis is the fifth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…Hugging FaceIntroducing the FFASR Leaderboard: Benchmarking ASR in the Real WorldHugging FaceAccelerating Transformers Fine-Tuning with NVIDIA NeMo AutoModelGitLabGitLab Patch Release: 19.1.1, 19.0.3, 18.11.6Google DeepMindIntroducing computer use in Gemini 3.5 FlashCohereWhen AI Doctors "See" What Isn't There: Why Better Accuracy Doesn't Mean Better VisionCohereAnanya SahuCohereMarzieh FadaeeStripeFour travel and hospitality trends from HITEC 2026More than 6,000 hospitality executives and operators gathered in San Antonio last week for the HITEC conference. The big topic: whether the industry’s AI investment is actually working. Across four days and over 50 meeti…SalesforceHow Agentforce Prevents Language Drift in 600K Daily Multilingual AI WorkflowsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Ishween Kaur, Senior Software Engineer on the Agentforce Agentic Reasoning team. As A…NetflixToward More Controllable AI Video Editing: An Early Research Exploration at NetflixBy Zhuoning Yuan , Ta-Ying Cheng , Benjamin Klein , Bahareh Azarnoush Introduction At Netflix, we build technology to help storytellers bring their creative visions to life and to help members discover the stories they l…MozillaPACT: Anonymous Credentials for the WebThis is the technical companion to our update on Distilled, “Keeping the web open and private in the bot era.” Here we take a deeper look at the problem space, the design we re proposing, and the problems still left to s…Mistral AIMistral OCR 4 : SOTA OCR for Document IntelligenceMistral AIOCR 4MetaHow Meta Engineered Ultra-Narrow Batteries for AI GlassesSmart glasses like the Ray-Ban Meta and Oakley Meta Vanguards need to pack enough energy to power features like cameras, speakers, AI workloads, and even a display. But it all has to fit into the glasses’ temple arms. So…Hugging FaceExperimenting with the proposed Cross-Origin Storage API in Transformers.jsHugging FaceShipping huggingface_hub every week with AI, open tools, and a human in the loopHugging FaceBuild real agentic apps using CUGA: two dozen working examples on a lightweight harnessGitHubI automated my job (and it made me a better leader)Explore how my day as a senior leader looks now that I use 40 automations to help, and learn more about some of my favorites. The post I automated my job (and it made me a better leader) appeared first on The GitHub Blog…GitHubGitHub joins coalition advocating for fixes to California AI Transparency Act to protect open sourceWe’re calling for targeted amendments to resolve conflicts with open source licensing and align with international transparency frameworks while preserving regulatory intent. The post GitHub joins coalition advocating fo…CohereWhy Cultural Awareness is Essential for Global AIAnthropicIntroducing Claude TagNetflixHow Netflix Simplified Batch Compute with KueueBy Alvin Bao , Alex Petrov , Jennifer Lai , Aidan Sherr , and Samartha Chandrashekar As a part of the journey to transition Netflix’s compute infrastructure to be more Kubernetes-native, we have leaned into incorporating…MicrosoftModels don’t have preferences, they have contextYou open a fresh chat, type What framework should I use for a web app? , and the model says React. You screenshot it, share it, and write Claude prefers React. It gets engagement. People nod along. A few reply with their…MicrosoftModels don’t have preferences, they have contextYou open a fresh chat, type What framework should I use for a web app? , and the model says React. You screenshot it, share it, and write Claude prefers React. It gets engagement. People nod along. A few reply with their…MetaAdopting AV1 for Real-Time Communication (RTC) at ScaleAdopting AV1 for real-time communication at Meta has been a multi-year effort spanning codec selection, device eligibility, rate control, and error resilience. We’re sharing the technical and operational challenges while…Hugging FaceWe got local models to triage the OpenClaw repo for FREE!*Hugging FacePP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M ParametersGrabScaling out Distroless adoption With AIDistroless adoption at Grab Grab is migrating from heavy base images to Distroless images to reduce security risks. By limiting each container to the application and its runtime dependencies, we shed non-essential binari…GitHubFrom pledge to practice: Building a more inclusive open source ecosystemLearn about the progress we’ve made toward our accessibility goals and how you can help make open source more inclusive. The post From pledge to practice: Building a more inclusive open source ecosystem appeared first on…AWSAWS Weekly Roundup: NY Summit recap, Local Zone in Hanoi, Grok 4.3 in Bedrock, price reductions, and more (June 22, 2026)Last week AWS Summit New York City brought together thousands of customers, partners, and builders for a free, one-day event showcasing the latest in cloud and AI innovation. Dr. Swami Sivasubramanian, VP of Agentic AI a…AWSRun isolated sandboxes with full lifecycle control: AWS Lambda introduces MicroVMsAWS launches a new serverless compute primitive, AWS Lambda MicroVMs. VM-level, isolated sandboxes with no shared kernel or resources between sessions. Rapid launch and resume, full lifecycle control, state preservation…GrabPalana (Part 2): Architecting isolation, identity, and auditability for AI agentsIntroduction In Part 1 , we introduced Palana , Grab’s Kubernetes-native secure execution platform for autonomous AI agents. We discussed the underlying need for isolated environments and covered its core design principl…NetflixData Projects: Managing Data Assets at Netflix ScaleBy Amer Hesson , Marcelo Mayworm , James Mulcahy , and Brittany Truong The Problem: Managing Assets at Netflix Scale Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands…NetflixThe Data Canary: How Netflix Validates Catalog MetadataBy Celina Amados At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated…NetflixPredicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planningby Emily Gill Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and bui…NetflixThinking Fast & Slow for a Personalized Notification Systemby Matthew Wood , Ishan Gupta , Kevin Mercurio, Devon Bryant , and Claire Dorman In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates a…NetflixVMAF v1: Good Is Not Good EnoughBy Christos G. Bampis , Zhi Li , Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? Wha…NetflixFrom Silos to Service Topology: Why Netflix Built a Real-Time Service MapBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez Nathan Fisher How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netflix runn…NetflixA Human-Augmenting Agentic Workflow for Causal InferenceBy Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus Introduction Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watchin…NetflixThinking Fast & Slow for a Personalized Notification Systemby Matthew Wood , Ishan Gupta , Kevin Mercurio, Devon Bryant , and Claire Dorman In his seminal book “Thinking, Fast and Slow,” Daniel Kahneman describes two systems that drive human cognition: System 1, which operates a…NetflixThe Evolution of Cassandra Data Movement at NetflixBy Guil Pires , Jennifer Prince , Jose Camacho , Ken Kurzweil , Phanindra Chunduru Background In a previous post, we introduced Data Bridge , a unified management plane for batch Data Movement at Netflix. Historically, s…NetflixPredicting Risk in Content Launches: How Data-Driven Insights can Transform Launch Planningby Emily Gill Each year, we bring the Analytics Engineering community together for an Analytics Summit — a multi-day internal conference to share analytical deliverables across Netflix, discuss analytic practice, and bui…NetflixData Projects: Managing Data Assets at Netflix ScaleBy Amer Hesson , Marcelo Mayworm , James Mulcahy , and Brittany Truong The Problem: Managing Assets at Netflix Scale Netflix’s Data Platform is vast. We have millions of tables in our data warehouse and tens of thousands…NetflixThe Data Canary: How Netflix Validates Catalog MetadataBy Celina Amados At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated…NetflixVMAF v1: Good Is Not Good EnoughBy Christos G. Bampis , Zhi Li , Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao Will this encode look good to Netflix members? Does switching to a new codec improve quality at the same bitrate and by how much? Wha…NetflixFrom Silos to Service Topology: Why Netflix Built a Real-Time Service MapBy Parth Jain , Rakesh Sukumar , Yingwu Zhao , Renzo Sanchez-Silva Nathan Fisher How we built a living map of our distributed infrastructure to help engineers understand dependencies, troubleshoot faster, and keep Netfli…NetflixA Human-Augmenting Agentic Workflow for Causal InferenceBy Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus Introduction Imagine asking a data agent to analyze the causal relationship between two variables, such as the effect of watchin…NetflixThe Evolution of Cassandra Data Movement at NetflixBy Guil Pires , Jennifer Prince , Jose Camacho , Ken Kurzweil , Phanindra Chunduru Background In a previous post, we introduced Data Bridge , a unified management plane for batch Data Movement at Netflix. Historically, s…GrabPalana (Part 1): Why Grab built a secure platform for autonomous AI AgentsAbstract Artificial intelligence (AI) agents are moving from experiments into everyday engineering workflows. They can read code, call application programming interfaces (APIs), run tests, create merge requests, answer S…GitHubHow we built an internal data analytics agentQubot, our internal Copilot-powered analytics agent, allows any GitHub employee to ask questions about our data in plain language. Here's what we learned as we built it. The post How we built an internal data analytics a…StripeWhat Link data tells us about AI spendingWe analyzed spending patterns across the 250 million customers paying with Link. We found that Link customers are spending more on AI than they were three months prior, investing heavily in platforms that let them build…Mistral AIFrontier AI LLMs, assistants, agents, servicesMicrosoftStop overloading your skillsYou built a skill for your technology. API references, authentication flows, SDK patterns, error handling, version info, all packed into one skill. The agent calls it, gets all that context, and generates code. The kicke…Hugging FaceIs it agentic enough? Benchmarking open models on your own toolingHugging FaceBeyond LoRA: Can you beat the most popular fine-tuning technique?Hugging FaceMosaicLeaks: Can your research agent keep a secret?GitLabAI Catalog updates for governance and operationsEnterprise AI adoption often stalls not because the technology isn t ready, but because admins can t answer the question their security team is asking: What s actually running in our environment, and who put it there? Gi…GitLabGitLab 19.1 releasedGitLabOne vulnerability view: From scanner coverage to AI governanceMost enterprises use a handful of different security scanners, each configured and enforced, project by project. With no single view of what scanners run where, policies drift, blind spots go undetected, and important pr…GitHubHow pull request limits are cutting down the noiseLearn how pull request limits can help manage contribution volume in your repositories, and see what’s next on the roadmap. The post How pull request limits are cutting down the noise appeared first on The GitHub Blog .DiscordHow to Manage Your Discord Desktop Notifications: A Complete GuideNot getting pinged for the conversations you wanna know about? This guide walks through server and channel notifications settings, Do Not Disturb, turning off specific notification sounds, and what to check when your pin…AWSAmazon ECS introduces new high-resolution metrics for faster service auto scalingAmazon Elastic Container Service (Amazon ECS) service auto scaling automatically adjusts task counts to meet workload demand with comprehensive scaling policies, including predictive scaling for recurring traffic pattern…AWSAnnouncing Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUsAnnouncing the general availability of Amazon Elastic Compute Cloud (Amazon EC2) G7 instances, delivering high performance GPU acceleration for AI inference, graphics, and data analytics workloads.AnthropicProject Fetch: Phase twoSalesforceMaintaining Code Quality at Agent Speed: 7 Patterns for Agentic EngineeringBy Amit Sharma and Antonio Garrote.How do you know whether code generated at agent speed can be trusted? That question is fast becoming one of the most important in software engineering and it deserves a serious answer.…MicrosoftWhen your agent extensions fight each otherThis is the fourth article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…Hugging FaceMolmoMotion: Language-guided 3D motion forecastingGitLabGitLab and Capgemini accelerate DevSecOps transformationWe re excited to share that GitLab and Capgemini have signed a global alliance partnership agreement. As a GitLab Select Partner, Capgemini will leverage the GitLab portfolio for clients globally, including GitLab Duo Ag…GitLabGitLab named a Leader in the 2026 Gartner® Magic Quadrant™ for DevSecOps PlatformsFor the fourth year running, Gartner has named GitLab a Leader in the 2026 Gartner® Magic Quadrant™ for DevSecOps Platforms . We believe this recognition reflects what our customers already see: The work of building soft…GitLabIntroducing the 2026 EMEA GitLab Partner Award winnersThe GitLab Partner Program continues to cultivate a thriving ecosystem of DevSecOps professionals committed to helping customers modernize their software development, from foundational DevSecOps practices to the latest A…GitHubGetting more from each token: How Copilot improves context handling and model routingHow GitHub Copilot is making more of each session go toward useful work, so your credits go further. The post Getting more from each token: How Copilot improves context handling and model routing appeared first on The Gi…CohereMusa TalluziCohereLLM Serving Fairness: No More Noisy NeighborsAWSIntroducing Amazon Bedrock Managed Knowledge Base for faster, more accurate enterprise AI applicationsAmazon Bedrock's new Fully Managed Knowledge Bases simplifies building enterprise RAG pipelines by providing native data connectors Smart Parsing for automatic multi-format data preparation, and an Agentic Retriever for…AWSTop announcements of the AWS Summit in New York, 2026A recap of the top announcements from AWS's New York Summit 2026AWSAWS DevOps Agent adds release management capabilities to assess code changes before production (preview)AWS DevOps Agent now offers release management capability in preview, reviewing code changes for release readiness and running autonomous release testing to help you ship code to production safely and with confidence.AWSAWS Security Agent adds threat modeling, Kiro power and Claude Code plugin, and moreAWS Security Agent now adds STRIDE-based threat modeling, full repo and PR code scanning with remediation across major Git platforms, and IDE integrations via Kiro power, Claude Code plugin, and MCP — letting developers…AWSProactively reduce tech debt autonomously with AWS Transform – continuous modernization (preview)AWS Transform – continuous modernization (preview) automatically scans code repositories to detect, prioritize, and remediate technical debt at scale.AWSAnnouncing Web Search on Amazon Bedrock AgentCore: Ground your AI agents in current, accurate web knowledgeAWS introduces Web Search on Amazon Bedrock AgentCore, a fully managed tool that enables agents to ground responses in current, cited web knowledge with zero data egress from customer's secured AWS environment. You can f…AnthropicAnthropic opens Seoul office and announces new partnerships across the Korean AI ecosystemSalesforceThe Agent Coding Maturity Curve: 9 Stages from Code Generation to Trusted AutomationMost developers begin with the same rush of excitement: the agent writes code, fixes bugs, explains unfamiliar systems, generates tests, and turns vague intent into something that looks runnable. For a moment, it feels l…MicrosoftCompeting against yourselfYou shipped a new CLI: better developer experience, modern architecture, and optimized for agents. You deprecated the old one, updated the docs, and blogged about it. Developers are migrating. Then someone asks an AI cod…GitHubWhat are git worktrees, and why should I use them?Git worktrees have been around since 2015, but it wasn't until recently they became popular. Learn what they are, how to use them, and why you might. The post What are git worktrees, and why should I use them? appeared f…Google DeepMindSecuring the future of AI agentsSecuring internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.Google DeepMindUnlocking UK house-building with AI-accelerated planningUK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.AWSAmazon S3 annotations: attach rich, queryable context directly to your objectsAmazon S3 now lets you attach up to 1 GB of rich, mutable, and queryable context directly to your objects using annotations, purpose-built for AI agents and autonomous workflows that need to discover, understand, and act…SalesforceHow Data 360 Segmentation Processes a Quadrillion Records Across Arbitrary Customer Data ModelsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Deepak Pushpakar, Software Engineering Architect for Segmentation and Activation with…PayPalWe have moved to https://developer.paypal.com/community/blogThe PayPal Developer Blog has a new home. You’ll continue to find the latest developer news, tutorials, product updates, and engineering insights here: https://developer.paypal.com/community/blog Thanks for being part of…GitHubAccelerating researchers and developers building multilingual AI with a new open datasetA new repository-level dataset, published on GitHub under CC0-1.0, helps researchers and developers discover multilingual developer content across READMEs, issues, and pull requests. The post Accelerating researchers and…GitHubGitHub Copilot CLI for Beginners: Overview of common slash commandsGitHub Copilot CLI for Beginners: Learn how to use slash commands to control your terminal AI agent. The post GitHub Copilot CLI for Beginners: Overview of common slash commands appeared first on The GitHub Blog .CohereCohere triples UK footprint with new London office to support R&D growthCohereA Community That Met Me HalfwayAWSAWS WAF adds AI traffic monetization capability to help content owners charge AI bots for content accessAWS WAF launches AI traffic monetization, a new Bot Control capability that enables content providers and publishers price, meter, and collect payment from AI bots and agents accessing their content and APIs. AWS WAF now…AWSAWS Weekly Roundup: AWS FinOps Agent in preview, Gemma 4 on Bedrock, Kiro Pro Max, and more (June 15, 2026)This week, New York City is hosting AWS Summit, bringing together builders, customers, and AWS teams for a full day of announcements, demos, and technical sessions at the Javits Center. I wrote blog posts for some of the…Microsoft ResearchIre identifies another LOTUSLITE specimenProject Ire examined a timely malware sample and determined its intent through reverse engineering—identifying LOTUSLITE characteristics even as most major EDR tools did not detect it. The post Ire identifies another LOT…Mistral AIFrontier AI LLMs, assistants, agents, servicesGitHubHow we made GitHub Copilot CLI more selective about delegationBetter orchestration, fewer handoffs, faster progress, without a single new knob. The post How we made GitHub Copilot CLI more selective about delegation appeared first on The GitHub Blog .DropboxHow Dropbox uses MCP and Dash to close the design-to-code security gapUsing an agentic AI system to surface threat models during code review and spot gaps between security requirements and implementation.AnthropicStatement on the US government directive to suspend access to Fable 5 and Mythos 5AnthropicTCS and Anthropic partner to bring Claude to regulated industriesAnthropicResults from the first Anthropic Public RecordStripeStripe Projects adds new agent integrations, more providers, and custom developer controlsOur data shows that agents are now fully capable of independently writing code and integrating with APIs like Stripe’s. And yet, many of the steps adjacent to writing code are still too hard for agents to do on their own…SlackAgentic Testing: Where Agents Fit in the E2E Testing StackAbstract Agent-driven end-to-end (E2E) tests add a new exploratory layer to testing, but should they replace traditional deterministic tests? We ran more than 200 agentic E2E workflows using the Playwright MCP, Playwrigh…SalesforceHow MuleSoft Is Raising the Trust Bar for AI-Generated CodeIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Melissa Cazalet, Senior Vice President of Software Engineering at MuleSoft, whose tea…MicrosoftYour agent just scaffolded a project from 2020Your agent ran a scaffold command. Project generated, dependencies resolved, no errors. Everything looks fine. Except it s based on the project structure from 2020, and neither you nor the agent noticed. How npx picks th…GitLabGitLab Patch Release: 19.0.2, 18.11.5, 18.10.8GitHubGitHub availability report: May 2026In May, we experienced nine incidents that resulted in degraded performance across GitHub services. The post GitHub availability report: May 2026 appeared first on The GitHub Blog .GitHubMaking secret scanning more trustworthy: Reducing false positives at scaleAlerts are more trustworthy and actionable when noise is reduced. See how we improved the verification step with context-aware LLM reasoning. The post Making secret scanning more trustworthy: Reducing false positives at…CohereZanele MunyikwaAnthropicDXC will integrate Claude into the systems banks, airlines, and other regulated industries rely onAnthropicDXC will integrate Claude into the systems banks, airlines, and other regulated industries rely onAnthropicIntroducing Claude CorpsSpotifyEncoding Your Domain Expert: The Context Layer Behind Spotify's Data AssistantAt Spotify, data problems used to follow a specific pattern. You'd look for the relevant dashboard, there... The post Encoding Your Domain Expert: The Context Layer Behind Spotify's Data Assistant appeared first on Spoti…MicrosoftSpec-Driven Development: A Spec-First Approach to AI-Native EngineeringAI has made software delivery faster, but speed alone does not guarantee better outcomes. As teams adopt AI-native development, the real challenge is keeping requirements, design, implementation, and validation aligned s…MicrosoftStop skillmaxxing, save your tokensYou built a dozen skills for your technology: authentication, CRUD, error handling, deployment, testing, monitoring. Then you installed a cloud platform bundle with 15 more covering diagnostics, storage, compliance, and…MicrosoftIs your agent extension actually working?This is the third article in a series about Agent Experience (AX): the practice of making AI coding agents work correctly with your technology. The series covers what you can and can t control in the agent stack, how to…LyftMetric Semantic Layer: How Lyft Governs and Scales Key Data DefinitionsWritten by Rohit Channe and Simran Mirchandani at Lyft. Motivation At Lyft, data isn’t just a resource — it’s woven into everything we do. Metrics drive key forecasts, steer operational decisions, and put our boldest hyp…LyftMetric Semantic Layer: How Lyft Governs and Scales Key Data DefinitionsWritten by Rohit Channe and Simran Mirchandani at Lyft. Motivation At Lyft, data isn’t just a resource — it’s woven into everything we do. Metrics drive key forecasts, steer operational decisions, and put our boldest hyp…GitLabGitLab on Google Cloud: Fully managed, compliant, and AI-readyYou can now run GitLab as a fully managed platform on Google Cloud, delivered by GitLab-certified managed service providers (MSPs) — with the latest Google AI models built in. Through MSP partners, including Beyond and D…GitLabGitLab: Built for the agentic engineering eraGitLab Transcend, our customer event showcasing our roadmap, success stories, and industry research recently wrapped. Here s what we announced and demonstrated: Next-generation source code management , a Git engine rebui…GitLabGitLab Flex: Commit once, reshape your seats and AI spendThe agentic era made your needs harder to predict, and the way you buy software hasn t caught up. Six months out, you don t know how many seats you ll need, how much AI your teams will consume, or which new capabilities…GitLabIntroducing GitLab Orbit: Full code and lifecycle context, in one queryAgents are good at writing code. They re far worse at navigating the system around it: the related code, the pipelines that run it, the deployments that ship it, the work items that asked for it, and the teams that own i…GitHubGive GitHub Copilot CLI real code intelligence with language serversInstall and configure LSP servers for GitHub Copilot CLI, replacing brute-force grep/decompile with real code intelligence. The post Give GitHub Copilot CLI real code intelligence with language servers appeared first on…DiscordGame On: Discord Is Backing the Next Generation of Dutch Gaming FoundersTogether with Techleap, we’re launching the Gaming Founders Circle: a hands-on program for the most ambitious gaming companies in the Netherlands to provide participants with a direct line to the people, investors, and n…DiscordUpdated Requirements to How Apps Access Data in ServersDiscord is updating requirements to how apps access certain data within servers, changing the review threshold and requiring annual review. Here s what s changing and why.Google DeepMindInvesting in multi-agent AI safety researchGoogle DeepMind and partners announce a $10M funding call for multi-agent safety research.Google DeepMindDiffusionGemma: 4x faster text generationCohereThe future of work debate has an evidence problemAWSNow available: Amazon EC2 M9g and M9gd instances powered by new AWS Graviton5 processorsAWS launches Amazon EC2 M9g and M9gd instances, powered by AWS Graviton5 processors. AWS Graviton5 is most powerful, and most energy efficient processor AWS has ever built, and offers up to 25% better compute performance co…SalesforceHow to Build Reliable AI Agents: 5 Engineering Patterns from a Production SystemBy Tuhin Kanti Sharma and Chirag Ramesh Hegde. If you ve built an AI agent that works perfectly in demos but becomes unpredictable in production, you ve probably already discovered that reliability is much harder than ca…LyftFrom Chaos to Clarity: How We Built a Unified, Self-Routing Support Ops Ticketing System at LyftWritten by Atul Gupta , Analytics Manager — LUS Support Ops, Lyft At Lyft, getting operators and riders connected quickly and reliably depends on more than technology — it depends on the teams working behind the scenes t…LyftFrom Chaos to Clarity: How We Built a Unified, Self-Routing Support Ops Ticketing System at LyftWritten by Atul Gupta , Analytics Manager — LUS Support Ops, Lyft At Lyft, getting operators and riders connected quickly and reliably depends on more than technology — it depends on the teams working behind the scenes t…GitLabMythos-class Claude Fable 5 arrives on GitLab Duo Agent PlatformUpdate - July 1, 2026: Fable 5 has been redeployed by Anthropic and is again available for customers on GitLab.com. Some routine coding and debugging tasks may temporarily route to Claude Opus 4.8 as Anthropic refines it…GitLabShai-Hulud copycat campaign targets Python developers through PyPI typosquattingGitLab s Vulnerability Research team has identified a coordinated supply chain attack on PyPI deploying a copy of the Shai-Hulud malware. We found five malicious packages: four typosquats impersonating Flask, Requests, a…GitHubFrom one-off prompts to workflows: How to use custom agents in GitHub Copilot CLICustom agents let GitHub Copilot CLI understand your stack and team workflows, turning one-off terminal prompts into repeatable, reviewable processes. The post From one-off prompts to workflows: How to use custom agents…DiscordHow We Moved Discord Voice to the EdgeMoving Discord’s voice and video onto Cloudflare s edge network. Closer servers, lower ping in most regions, and a few real bugs getting there.Google DeepMindPowering the future of robotics in EuropeGoogle DeepMindIntroducing Gemma 4 12B: a unified, encoder-free multimodal modelGoogle DeepMindFluid, natural voice translation with Gemini 3.5 Live TranslateGemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.CohereNorth Mini Code: Agentic Coding for DevelopersAWSAnthropic Claude Fable 5 on AWS: Mythos-class capabilities with built-in safeguards now availableAWS announces the availability of Claude Fable 5 on Amazon Bedrock and Claude Platform on AWS. Claude Fable 5 delivers Mythos-level capabilities available to all customers, with strong safeguards designed to make it safe…AnthropicClaude Fable 5 and Claude Mythos 5AirbnbScaling beyond one: How Airbnb evolved its data architecture for a multi-product worldHow Airbnb’s data engineers and analytics engineers built a consistent and flexible data modeling framework to support the expansion into Homes, Experiences, and Services. By : Patrick Lam , Namrata Lamba , Jamie Stober…SalesforceScaling Zero Copy from 1 Trillion to 120 Trillion Rows with File FederationIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Srini Krishnamoorthy, Vice President of Engineering for Data 360. Srini leads the evo…MicrosoftMicrosoft Build 2026 recap: vision, launches, and top sessionsCatch up on Microsoft Build 2026 with the vision lead-off, top developer announcements, and must-watch sessions across the Microsoft developer ecosystem. The post Microsoft Build 2026 recap: vision, launches, and top ses…GitHubGitHub for Beginners: Answers to some common questionsFind the answers to some of the most common GitHub-related questions. The post GitHub for Beginners: Answers to some common questions appeared first on The GitHub Blog .DiscordIntroducing: You BarIntroducing the You Bar: an update to the Discord mobile app that celebrates your identity, simpler navigation, and a peek at what’s next.Google DeepMindMeasuring the impact of learning with AI in Sierra Leone and beyondResults from a randomized controlled trial show the potential of Gemini’s Guided Learning feature to boost engagement and accelerate learning.AWSAWS Weekly Roundup: BYOM for Amazon RDS for SQL Server, AWS IoT Device SDK for Swift, and more (June 8, 2026)This week, the AWS IoT Device SDK for Swift reached general availability. As a member of the Swift Server Workgroup (SSWG), this one caught my attention. The SDK brings production-ready MQTT 5 connectivity, Device Shadow…CohereAI for DevelopersCohereCompany NewsCohereProduct LaunchCohereEnterprise AICohereResearchAWSTry the new console experience in Amazon Bedrock, optimized for Anthropic- and OpenAI-compatible APIsYou can use the new console experience on Amazon Bedrock to browse and compare the latest AI models side by side, organize work into projects with streamlined evaluation workflows, and access project-aware live documenta…StripeRethinking risk in the age of AIJoin senior risk and payments leaders in Seattle to explore how AI is reshaping fraud strategy. Seats are limited.StripeThe future of agentic commerce is hereExplore how AI agents are transforming commerce at Stripe’s Agentic Commerce Next roadshow. Reserve your spot in Seattle.StripeNew ways to turn global demand into revenueAt Sessions 2026, Stripe unveiled dozens of products and capabilities to help businesses turn global demand into revenue. See how to go global faster with localized checkout and Adaptive Pricing, smarter fraud tools, mul…SalesforceHow Engineering 360 Unified Operations at Scale and Reached 80% AdoptionBy Shiva Nimmagadda, Arun Lakshmi Narayanan, and Arun Gangavarapu. Salesforce engineering teams encountered a significant operational hurdle as the organization scaled. Critical data lived across dozens of fragmented das…GitHubGitHub Universe is back: All together now, in the agentic eraGitHub Universe is back: returning to the historic Fort Mason Center in San Francisco on October 28–29, 2026. The post GitHub Universe is back: All together now, in the agentic era appeared first on The GitHub Blog .DiscordDiscord Patch Notes: June 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.CohereTalking to a 4-Year-Old: A Multilingual Benchmark for Children's AI CompanionsAirbnbSitar-agent: Building a reliable dynamic configuration sidecar at scaleHow Airbnb built a Kubernetes sidecar to deliver dynamic configuration reliably at scale. By : Bo Teng , Cosmo Qiu , Siyuan Zhou , Ankur Soni , Xin Huang , Willis Harvey Introduction In our previous post , we explored Ai…StripeHelping businesses optimize network costs with the Visa Digital Commerce Authentication Program (DCAP)We moved quickly to help Stripe businesses take advantage of DCAP and capture interchange savings while protecting authorization rates. Here’s what we did.SpotifyCoding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at SpotifyAt Code with Claude, Spotify’s chief architect shared how we make both teams and AI agents more effective. The post Coding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at Spotify appeared…SalesforceHow Agentforce Conversation Client Accelerated Accessibility Remediation by 5x Using AI-Driven WorkflowsBy Prasanna Krishna Sanagala, Ronak Shah, Sandeep Tailor, and Mani Manjari Velnati. In our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight P…NetflixDynamic Repartitioning for Time Series WorkloadsBy Rajiv Shringi , Kaidan Fullerton , Oleksii Tkachuk and Kartik Sathyanarayanan Introduction Netflix’s TimeSeries Abstraction is a scalable system for ingesting and querying petabytes of temporal event data with millise…NetflixDynamic Repartitioning for Time Series WorkloadsBy Rajiv Shringi , Kaidan Fullerton , Oleksii Tkachuk and Kartik Sathyanarayanan Introduction Netflix’s TimeSeries Abstraction is a scalable system for ingesting and querying petabytes of temporal event data with millise…MetaLights Out, Systems On: Validating Instant Power Loss ReadinessWe’re introducing Instantaneous PowerLoss Storm, a new testing paradigm within Meta’s infrastructure for handling and mitigating instant or zero-notice power loss in our data centers. We’re sharing: how we built readines…CohereCo/plot: Supporting the research process through visualizationAWSImprove your application resilience with Amazon Cognito multi-Region replicationAmazon Cognito now offers multi-Region replication that automatically synchronizes user data, credentials, and pool configurations to a secondary AWS Region, enabling uninterrupted authentication during regional failover…AnthropicWhat we learned mapping a year’s worth of AI-enabled cyber threatsAnthropicIntroducing the Services Track and Partner Hub of the Claude Partner NetworkAnthropicWhat we learned mapping a year’s worth of AI-enabled cyber threatsInstacartSemantic IDs: Product Understanding at ScaleKey Contributors: Shrikar Archak, Karuna Ahuja, Soroush Sobhkhiz, Marko Avdalovic, Xiyu Wang, JiChao Zhang, Hao Yan, Chris Hartley Introduction Operating a grocery catalog at Instacart’s scale means managing millions of…InstacartFrom Scoring to Spelling: Rebuilding Ads Retrieval at InstacartKey Contributors: Karuna Ahuja, Marko Avdalovic, Soroush Sobhkhiz, Shrikar Archak, Xiyu Wang, Ji Chao Zhang, Hao Yan Introduction Every time a user opens Instacart, they see product recommendations: on the retailer home…InstacartSemantic IDs: Product Understanding at ScaleKey Contributors: Shrikar Archak, Karuna Ahuja, Soroush Sobhkhiz, Marko Avdalovic, Xiyu Wang, JiChao Zhang, Hao Yan, Chris Hartley Introduction Operating a grocery catalog at Instacart’s scale means managing millions of…InstacartFrom Scoring to Spelling: Rebuilding Ads Retrieval at InstacartKey Contributors: Karuna Ahuja, Marko Avdalovic, Soroush Sobhkhiz, Shrikar Archak, Xiyu Wang, Ji Chao Zhang, Hao Yan Introduction Every time a user opens Instacart, they see product recommendations: on the retailer home…GitHubGitHub Copilot app: The agent-native desktop experienceAt Microsoft Build 2026, GitHub introduced new tools, updates, and surfaces so agents can work the way you already work. The post GitHub Copilot app: The agent-native desktop experience appeared first on The GitHub Blog…AnthropicExpanding Project GlasswingAirbnbWhen history fails you, borrow from geographyHow Airbnb used sequential geographic recovery signals and prior propagation to generate reliable corridor-level forecasts when local data was scarce. By: Harrison Katz The problem with unprecedented shocks Almost every…CohereRWS and Cohere Build Enterprise AI TranslationAWSGet started with OpenAI GPT-5.5, GPT-5.4 models, and Codex on Amazon BedrockOpenAI frontier models GPT-5.5 and GPT-5.4, and Codex, the OpenAI coding agent, are available on Amazon Bedrock. Deploy frontier models on Bedrock's high performance inference engine with built-in security, governance, a…AWSAWS Weekly Roundup: Claude Opus 4.8 on AWS, Aurora MySQL with Kiro Powers, and more (June 1, 2026)In my last Week in Review post, I shared what I’d been hearing from customers in the AI-Driven Development Lifecycle (AI-DLC) workshops I’ve been delivering. Last week I was back at it, this time in Denver for a two-day…AnthropicAnthropic confidentially submits draft S-1 to the SECNetflixHigh-Throughput Graph Abstraction at Netflix: Part IBy Oleksii Tkachuk , Kartik Sathyanarayanan , Rajiv Shringi Introduction Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These…NetflixHigh-Throughput Graph Abstraction at Netflix: Part IBy Oleksii Tkachuk , Kartik Sathyanarayanan , Rajiv Shringi Introduction Netflix has a diverse range of graph use cases, each serving specific business needs with unique functionality and performance requirements. These…GrabFrom decentralized Docs-as-Code to a centralized repository: Evolving Grab's documentation strategyIntroduction: The journey of documentation at Grab In early 2021, Grab adopted a Docs-as-Code approach to address gaps in our technical documentation processes, as illustrated in our blog post Embracing a Docs-as-Code .…StripeSolo founding is at an all-time high: Top performers have these traits in commonIn 2025, solo founders in the top decile generated 61 times the revenue of the median solo founder in their first six months. We analyzed the data to understand what drives that gap.SlackSlack AI: The Path to Multi-CloudIn early 2023, Slack faced a foundational challenge: serving Large Language Models (LLMs) at enterprise scale with the security, reliability, and performance our customers expect. Over three years, we evolved from basic…Microsoft ResearchData Formulator 0.7: AI-powered data analytics for enterprise dataData Formulator introduces AI-powered analytics for enterprise data workflows. Data teams can easily bring enterprise data into an AI-ready workspace where users can explore, analyze, and visualize data with AI agents to…Mistral AIAI Now Summit 2026Mistral AIVibe gets to work.Mistral AIIntroducing Search ToolkitMicrosoftImprove your agentic developer tools by grounding in Microsoft LearnDevelopment workflows span terminals, IDEs, background agents, and custom assistants. What matters is whether they draw from the same current source. Learn MCP Server gives any MCP-compatible agent direct access to curre…GitLabGitLab Patch Release: 19.0.1, 18.11.4, 18.10.7GitLabClaude Opus 4.8 on GitLab: Complex agentic work, less disruptionAnthropic s latest model on GitLab is built for precise execution across complex multi-step agent work. Agents fail most often on complex, multi-step work: tasks that span multiple tools and go from intent to production…GitLabAgentic coding is only as good as its contextEvery week, another coding agent demo shows a prompt turning into a merge request in under five minutes. These demos often highlight a narrow use case not yet in production, and they skip everything that happens after th…GitHubStill a developer. Just outside. Our latest GitHub Shop collection is here.The ESC collection lets you escape the confines of your desk and get out into the sun where good ideas are bound to happen. The post Still a developer. Just outside. Our latest GitHub Shop collection is here. appeared fi…DropboxBeyond code generation: rethinking engineering productivity in the age of AI agentsHow Dropbox is moving from AI tools that assist engineers to agentic systems that can execute scoped tasks, and how we’re building platforms to support those workflows.DiscordOfficial Discord Integrations for Steal a Brainrot, Grow a Garden, Brookhaven RP, and moreHow some of the most popular Roblox games have integrated Discord account linking to enhance social features and safety capabilities for official community servers.CohereWhat is Model Context Protocol? A practical guide to MCPCohereThe Enterprise Guide to AI in Business IntelligenceCohereAI Governance Challenges: How to Scale ResponsiblyAWSIntroducing the next generation of Amazon OpenSearch Serverless for building your agentic AI applicationsAWS rebuilt Amazon OpenSearch Serverless from the ground up for agentic AI and dynamic workloads. Get instant autoscaling and up to 60% cost savings.AWSIntroducing the next generation of AWS Resilience Hub for generative AI-based SRE resilience journeyAWS launches the next generation of AWS Resilience Hub with a significantly expanded experience that brings together a new application model, dependency discovery assessment, generative AI-powered failure mode analysis,…AnthropicIntroducing Claude Opus 4.8AnthropicAnthropic raises $65B in Series H funding at $965B post-money valuationAnthropicIntroducing Claude Opus 4.8YelpBeyond the Menu Tree: How Yelp Built a Smarter Customer Success Chatbot with AIThe Evolution of Support: From Fixed Phrases to Conversation At Yelp, delivering responsive and accurate customer support is a core priority. For years, our legacy Customer Success (CS) Chatbot provided support by guidin…StripeExpanding Stripe Radar to protect more of your businessRadar now blocks high-risk transactions across all supported payment methods; defends against new fraud types like multi-account abuse and pay-as-you-go abuse, regardless of which payment processor you use; and gives pla…SalesforceAgentforce’s Agent Script: Building Deterministic Control for Enterprise AI WorkflowsIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Elijah Ben Izzy, Software Engineering Architect at Salesforce. Elijah is building Age…Microsoft ResearchExtending Human Intelligence Through AIUnderstanding AI as an extension of human intelligence—not a replacement for it—offers a more grounded path for building trustworthy AI systems. The post Extending Human Intelligence Through AI appeared first on Microsof…Mistral AIPhysics AI research that’s shaping the industry.Mistral AIIntroducing physics AI at Mistral: the foundation for engineering acceleration.MicrosoftHow AI coding agents actually use your technologyYou ship an SDK, a CLI, an API, and developers use it. Now AI coding agents use it too, except they use it differently than humans do. Most of the time you have no idea what s actually happening between developer types a…CohereCohere and Mila Partner on Quebec French AIAWSMeet Our Newest AWS Heroes – May 2026We’re excited to welcome four outstanding community leaders as our newest AWS Heroes. These individuals embody the spirit of collaboration and knowledge sharing that makes the AWS community thrive. From building AI-power…AnthropicAnthropic opens Milan office to support Italian enterprise, research, and developersAnthropicAnthropic opens Milan office to support Italian enterprise, research, and developersSalesforceBuilding an Enterprise Agent Platform: Enforcing Identity, Data, and API GovernanceWhile enterprises deploy AI agents at a rapid pace, their governance strategies often remain fragmented. Most organizations enforce identity, data access, and API security in separate silos, which creates dangerous gaps…MetaSilverTorch: Index as Model — A New Retrieval Paradigm for Recommendation SystemsWe’re introducing SilverTorch, a reimagining of recommendation systems that unifies all retrieval components for user generated content under a unified architecture. SilverTorch shows up to 23.7x higher throughput compar…GitLabFull security scanner coverage of your codebase in minutesAcross the industry, every CI/CD platform faces the same challenge: As organizations grow, manually configuring scanners to run across every pipeline definition file isn t scalable. AI is accelerating how fast teams ship…GitLabReduce supply chain risk with SBOM-based dependency scanningThird-party code dominates most codebases, and four recent supply chain incidents show how a single compromised package can ripple into every project that depends on it. AI is compounding this problem: Research suggests…EtsyShaping Product Understanding with Contrastive Reinforcement LearningEtsy’s marketplace is defined by the creativity and craftsmanship of our sellers and the hundreds of millions of highly diverse products they offer. You can find silversmiths who cold-forge recycled sterling silver, weav…CohereOur 2026 Summer Merch Collection Is HereAnthropicAnthropic appoints KiYoung Choi as Representative Director of Korea ahead of Seoul office openingAWSAWS Weekly Roundup: AWS Local Zones in Istanbul, open-source ExtendDB, Kiro Web, and more (May 25, 2026)There’s something genuinely energizing about working with startups — something I’ve been doing intensely for more than two years now. Startups operate at a different frequency: the urgency is real, the constraints are ti…AnthropicAnthropic co-founder Chris Olah's remarks on Pope Leo XIV's encyclical "Magnifica humanitas"Mistral AIEmmi joins Mistral to accelerate the AI-native industrySalesforceAgent Fabric Context Catalog and the Future of AI GovernanceModern agents no longer execute within predictable application boundaries. They invoke APIs dynamically, retrieve enterprise context through MCP servers, orchestrate workflows across multiple platforms and interact with…Mistral AIConnect the dots: Build with built-in and custom MCPs in StudioMistral AIRemote agents in Vibe. Powered by Mistral Medium 3.5.Mistral AIConnect the dots: Build with built-in and custom MCPs in StudioMistral AIRemote agents in Vibe. Powered by Mistral Medium 3.5.InstacartHow AI Changes the Role of Applied ScientistsLevi Boxell, Tilman Drerup, Alexandr Lenk The Economics Team at Instacart is an applied science team that operates at the intersection of machine learning engineering and economics. Similar to other applied science teams…InstacartHow AI Changes the Role of Applied ScientistsLevi Boxell, Tilman Drerup, Alexandr Lenk The Economics Team at Instacart is an applied science team that operates at the intersection of machine learning engineering and economics. Similar to other applied science teams…GrabThe Hugo evolution: Engineering Grab's unified, one-click data ingestion platform with Apache FlinkIntroduction Data drives every decision we make at Grab. As our operations scale, so does our need for robust, real-time data ingestion and processing frameworks. Enter Hugo: our self-service data platform that has long…YelpHow Partition Access Visualizations Reduced our Data Lake S3 Cost by 33%Introduction In large analytics environments, data teams often struggle to answer deceptively simple questions, like who their stakeholders are and how their data is being used. At Yelp, we address this by visualizing ac…PinterestMaking User-Sequence Data More Cost-Efficient, Faster, and Easier to UseAuthors ( listed alphabetically ) Ads Feature Engineering Infra team: Ajay Venkatakrishnan, Le Zhang Core ML Infra team: Eric Shang, Pihui Wei ML Data team: Connor Votroubek, Yi He User Understanding team: Camilo Munoz,…Microsoft ResearchVega: Zero-knowledge proofs for digital identity in the age of AIVega turns a full credential into a single proof, sharing only what is needed and nothing more, with performance that works in real apps. The post Vega: Zero-knowledge proofs for digital identity in the age of AI appeare…Microsoft ResearchMagenticLite, MagenticBrain, Fara1.5: An agentic experience optimized for small modelsMagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to support efficient agentic performance on ev…MozillaAnnouncing Web Serial Support in FirefoxSupport for Web Serial in Firefox 151 for Desktop Firefox can now connect directly to microcontrollers, development boards, 3D printers, power meters, and other serial-connected hardware from the web. Starting in Firefox…MicrosoftThe AX stack: what’s fixed, where you can winAI coding agents promise to make you more productive. On the surface they do, but in practice they fall short: agents generate code that doesn t compile, use a deprecated SDK, or pick the wrong service entirely. Is it yo…GitLabTrack CI component usage across your organizationIf your platform team publishes standardized pipeline components, you ve probably encountered this: once they re out in the wild, you lose visibility. You can t see if anyone’s actually using it, who s on which version,…GitLabGitLab 19.0 releasedGitLabMore AI models for GitLab Duo Agent Platform Self-HostedCustomers running GitLab Duo Agent Platform Self-Hosted operate under constraints many software teams don t face: data residency mandates, air-gapped networks, and compliance regulations that prohibit sending source code…GitLabManage CI/CD credentials with GitLab Secrets ManagerMany credential leaks start with a developer who needs a credential, doesn’t have a good place to put it, and improvises. It lands in an over-scoped CI/CD variable, a config file, or a .env committed “just for a moment.”…GitLabTransform MRs from manual tasks to an automated workflowAI made writing code dramatically faster, but the work between opening a merge request and merging it has stayed almost entirely manual. Assigning reviewers, addressing feedback round after round, untangling conflicts, r…DropboxIntroducing Nova, our internal platform for coding agentsNova lets engineers run multiple coding sessions in parallel and lets internal systems use AI agents as part of automated workflows.DiscordMaking It Easier Than Ever to Connect with Friends in League & VAL!Link your Discord and Riot accounts to sync your friends lists, show your in-game activity as your Discord status, and invite your Discord friends directly to your League or VAL lobby.Google DeepMindWe’re launching the Google DeepMind Accelerator program in Asia Pacific to tackle environmental risksYelpOptimizing Our Build Times by Migrating from Webpack to RspackOver the years, Webpack has remained the bundler of choice for many JS projects, including here at Yelp. While it has served us well, its speed has increasingly become a bottleneck as our monorepo continues to grow. Fort…CohereIntroducing Command A+: Making sovereign agentic capabilities available to allCohereCohere Releases Command A+: An Open-Source Enterprise AI Model Built for Sovereign Critical InfrastructureCohereAnnouncing strategic MOUs with Indra Group and Multiverse ComputingSalesforceHow Salesforce Built an AI Security Agent for Autonomous Threat TriageIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Mor Levi, Vice President of Detection, Analysis and Response at Salesforce, who leads…MicrosoftAgentic-Agile: Why Agent Development Needs Agile (Not Just Prompts)A bad system will beat a good person [or agent] every time ~Dr. William Edwards Deming (with apologies) I started vibe coding by writing prompts (often dictated into my phone), refining them with an agent in M365 Copilot…CohereCohere acquires Reliant AI to expand sovereign enterprise AI for the global biopharma and healthcare sectorsAnthropicKPMG integrates Claude across its core business and workforce of more than 276,000 in strategic allianceAnthropicWidening the conversation on frontier AIAnthropicWidening the conversation on frontier AIAirbnbScaling Airbnb’s identity graph with a unified knowledge graph infrastructureHow Airbnb shifts from PaaS to an internal knowledge graph infrastructure at scale. By: Lucen Zhao , Shukun Yang , Ashish Jain Knowledge graphs offer a natural and powerful way to represent relationships between entities…SpotifyBetter Experiments with LLM Evals — A funnel, not a forkTL;DR LLM evals, automated judges that assess relevance, coherence, and quality at scale, are a powerful new... The post Better Experiments with LLM Evals — A funnel, not a fork appeared first on Spotify Engineering .GitLabBeyond BYOK: Why governance matters for AI agentsGitHub recently announced that Copilot CLI now supports bring-your-own-key (BYOK) and locally running models. Developers can route CLI requests through their own model provider or run a local model entirely offline. But…GitLabCodex and GitLab: From code fix to productionCodex, a coding agent, is a lot of fun when you are deep in the terminal. Point it at a repository, give it a focused task, and it gets to work fast. It reads the code, proposes a fix, runs commands, and helps you move f…DiscordEvery Voice and Video Call on Discord Is Now End-to-End EncryptedAs of March 2026, E2EE is now enforced for every voice and video call on Discord. This represents a multi-year commitment, and Discord’s VP of Engineering is here to talk about why it matters.Google DeepMindFast-tracking genetic leads to reverse cellular agingBiologists use Co-Scientist to find novel factors that successfully rejuvenate human cells.AWSAWS Weekly Roundup: AWS Transform at 1 year, Claude Platform on AWS, EC2 M3 Ultra Mac instances, and more (May 18, 2026)Just a year ago, we launched AWS Transform for .NET, Mainframe and VMware workloads, the first agentic AI service purpose-built for modernizing enterprise applications at scale. At re:Invent 2025, we introduced AWS Trans…AnthropicAnthropic acquires StainlessGoogle DeepMindMaking it easier to understand how content was created and editedWe're expanding our tools to help you understand how content was created and edited across the web.Google DeepMindGemini for Science: AI experiments and tools for a new era of discoveryA collection of science tools and experiments to expand the scale and precision of scientific exploration.Google DeepMindIntroducing Google Antigravity 2.0Google DeepMindIntroducing Gemini OmniGoogle DeepMindSimulate real-world places with Project Genie and Street ViewWe’re expanding access to Google AI Ultra subscribers globally and introducing a new capability powered by Street View.Google DeepMindHow WeatherNext helped the National Hurricane Center better predict Hurricane Melissa’s historic landfall in JamaicaLearn how our WeatherNext AI model help forecasters give communities unprecedented time to prepare ahead of the historic Hurricane Melissa.Google DeepMindUncovering repurposed medicines to fight liver fibrosisStanford geneticist uses Co-Scientist to help find new treatments for chronic liver disease and liver fibrosis.Google DeepMindUniting biological toolkits for a new approach to ALSCo-Scientist unites Boston Children’s Hospital and MIT’s labs to explore new RNA-based treatments for ALS.Google DeepMindAccelerating discovery of liver disease mechanismsFilippo Menolascina uses Co-Scientist to identify new liver disease treatments and explain why existing drugs only help certain patients.Google DeepMindOpening new paths in aging researchCalico Life Sciences uses Co-Scientist to connect scattered findings and generate new leads in aging research.Google DeepMindFinding the molecular switches behind new infectious diseasesClare Bryant uses Co-Scientist to identify genetic triggers in emerging infectious diseases.Google DeepMindStrengthening Singapore’s AI Future: A New National PartnershipGoogle DeepMind and Singapore partner to apply frontier AI to address complex challenges across health, education, and sustainability and more.Microsoft ResearchFurther Notes on Our Recent Research on AI Delegation and Long-Horizon ReliabilityOur recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows. We appreciate the interest in this work and want to clarify several…GrabHow AI is transforming analytics at GrabIntroduction At Grab, analytics sits close to almost every decision that matters. Our north star is the democratization of intelligence, ensuring that anyone making a business call has immediate access to trustworthy ans…GrabScaling developer experience: How we improved Android Studio in a large monorepoIntroduction Long integrated development environment (IDE) sync/indexing times can quietly erode developer productivity, making code navigation sluggish, spiking memory usage, and slowing down Jetpack Compose preview upd…Google DeepMindGemini 3.5: frontier intelligence with actionGemini 3.5 is built to help you execute complex, agentic workflows.SalesforceCreating a Multi-Tenant AI Agent Platform Handling 7K+ Sessions Without Cross-Team InterferenceIn our Engineering Energizers Q A series, we highlight the engineering minds driving innovation across Salesforce. Today, we spotlight Priyanka Saraf, Senior Software Engineer on the Agentforce Foundations team. Priyanka…InstacartScaling Personalized Marketing for Multi-Tenant Commerce PlatformsTL;DR Background: Marketing Across Marketplace and Storefront Instacart operates across two distinct commerce experiences: Instacart Marketplace, our first-party consumer marketplace Storefront Pro, our white-label e-com…InstacartScaling Personalized Marketing for Multi-Tenant Commerce PlatformsTL;DR Background: Marketing Across Marketplace and Storefront Instacart operates across two distinct commerce experiences: Instacart Marketplace, our first-party consumer marketplace Storefront Pro, our white-label e-com…AWSAmazon Bedrock introduces new advanced prompt optimization and migration toolAmazon Bedrock Advanced Prompt Optimization enables customers to optimize their prompts for their current model or migrate prompts to new models faster than before with built-in evaluation feedback loops. Optimize your p…AnthropicAnthropic partners with the Gates FoundationAnthropicPwC is deploying Claude to build technology, execute deals, and reinvent enterprise functions for clientsMicrosoft Researchmimalloc: A new, high-performance, scalable memory allocator for the modern eramimalloc is an open-source, modern, scalable memory allocator that is a drop-in replacement for malloc and free. It is relatively small (~12K lines), with clear internal data structures, and is easy to build and integrat…Microsoft ResearchGridSFM: A new, small foundation model for the electric gridIntroducing GridSFM, a small foundation model that can predict AC optimal power flow in milliseconds, boosting efficiency and unlocking cost savings. Learn how GridSFM gives grid operators direct visibility into congesti…MetaReel Friends: Building Social Discovery that Scales to BillionsOn its face the new Friend Bubbles feature looks simple enough. It highlights Reels your friends have watched and reacted to. But sometimes the features that seem the most straightforward require the deepest engineering…DiscordCelebrate Discord’s 11th Birthday with an Exclusive Set of Emoji and WallpapersDiscord is turning 11 this year! To celebrate, we made over twenty Discord-themed emojis, along with over twenty wallpapers, and even a digital poster for everyone to download, for free!AnthropicIntroducing Claude for Small BusinessAirbnbViaduct 1.0 and the future of Airbnb’s data meshMoving from an internal tool to a community-driven, production-ready data mesh. By : Ryan Tanner , Raymie Stata , Adam Miskiewicz Introduction We’re excited to announce the 1.0 release of the Viaduct. This release marks…PinterestAn Engineer’s Guide to Better AI Skills: Implementing a Testing Process to Optimize Agent…An Engineer’s Guide to Better AI Skills: Implementing a Testing Process to Optimize Agent Performance in Any Repository or Skill Author: Daniel Reed The tech industry is currently seeing a massive overhaul in the way we…Microsoft ResearchAdvancing AI for materials with MatterSim: experimental synthesis, faster simulation, and multi-task modelsMatterSim is expanding what AI can do for materials science—from faster large-scale simulations to MatterSim-MT, a new multi-task model for simulating properties beyond potential energy surfaces alone. The post Advancing…MetaMigrating Data Ingestion Systems at Meta ScaleMeta’s data ingestion system, which our engineering teams leverage for up-to-date snapshots of the social graph, has recently undergone a significant revamp to enhance its reliability at scale. Moving from our legacy sys…Google DeepMindCo-Scientist: A multi-agent AI partner to accelerate researchIntroducing Co-Scientist, a collaborative AI partner built with Gemini to help researchers accelerate scientific breakthroughs.AWSAmazon Redshift introduces AWS Graviton-based RG instances with an integrated data lake query engineAmazon Redshift RG instances, powered by AWS Graviton, run data warehouse and data lake workloads up to 2.4x as fast as RA3 instances at 30% lower price per vCPU. Its integrated data lake query engine supports open table…StripeFive vertical SaaS insights from Sessions 2026AI is forcing platforms to expand beyond pure software. See how vertical SaaS platforms are using payments, financial services, and agentic commerce to build more durable businesses.Microsoft ResearchSocialReasoning-Bench: Measuring whether AI agents act in users’ best interestsUsing SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instructions to optimize for user interest. The…MetaLabyrinth 1.1: Making End-to-End Encrypted Backups Even More ReliableWe’re rolling out version 1.1 of Labyrinth, the encrypted storage system and protocol that secures messages and history on Messenger. Labyrinth 1.1 enhances the reliability of end-to-end encrypted backups with a new sub-…DiscordHow to Use Nitro: A Beginner’s Guide to Discord’s Premium SubscriptionWhat’s Discord Nitro all about? What perks does it give, and how can you get it? If you’re looking to expand your Nitro knowledge, you’re in the right place.DiscordNitro Now Comes with Xbox Game Pass and New Benefits. Welcome to Nitro Rewards.As we hit Nitro’s 10-year anniversary, we re launching Nitro Rewards: a brand-new benefits program built with some of the biggest names in gaming. See what’s coming for Nitro members, for no added cost.AWSAWS Weekly Roundup: Amazon Bedrock AgentCore payments, Agent Toolkit for AWS, and more (May 11, 2026)My most exciting news of last week: Amazon Bedrock AgentCore previewed the first managed payment capabilities enabling AI agents to autonomously access and pay for APIs, MCP servers, web content, and other agents. Built…PinterestEnhancing Ad Relevance: Integrating Real-Time Context into Sequential Recommender ModelsHuiqin Xin | Machine Learning Engineer II, Ads Vertical Modeling; Lakshmi Manoharan | Senior Machine Learning Engineer, Ads Vertical Modeling; Karthik Jayasurya | Staff Machine Learning Engineer, Ads Signals; Ziwei Guo |…NetflixScaling ArchUnit with Nebula ArchRulesBy John Burns and Emily Yuan Introduction At Netflix, we operate using a polyrepo strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repo…NetflixScaling ArchUnit with Nebula ArchRulesBy John Burns and Emily Yuan Introduction At Netflix, we operate using a polyrepo strategy with tens of thousands of Java repositories. This means that we need to have ways of sharing common build logic across these repo…DiscordHow Discord Automates ScyllaDB Clusters at ScaleYou ve been asked to stand up a brand-new database cluster, meaning a whole day of configuring dozens of nodes, validating replication, wiring up dual-write pipelines… what if this whole ordeal took less than two hours?…BAIR (Berkeley)Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling.apr-fig { text-align: center; margin: 1.35em 0; line-height: 1.4; } .apr-fig--wide img { display: inline-block; width: 100%; max-width: 100%; height: auto; vertical-align: middle; } .apr-fig--wide-0-8 { max-width: 80%;…MozillaBehind the Scenes Hardening Firefox with Claude Mythos PreviewTwo weeks ago we announced that we had identified and fixed an unprecedented number of latent security bugs in Firefox with the help of Claude Mythos Preview and other AI models. In this post, we’ll go into more detail a…GrabEnhancing Flink deployment with shadow testingIntroduction Ensuring the reliability of Apache Flink deployments in Grab is crucial for the availability of our business-critical, real-time applications. While all applications are tested in a staging environment befor…DiscordStock Up in the New Rust Shop! Enjoy a Discord-Only 20% Sale on Most Items until 5/21Starting today, you can now browse, purchase, and even gift and wishlist in-game items for Rust directly in Discord! See how it all works, and learn about a hefty two-week Discord-exclusive launch discount on a wide sele…Google DeepMindAlphaEvolve: How our Gemini-powered coding agent is scaling impact across fieldsExplore how AlphaEvolve's Gemini-powered algorithms are driving impact across business, infrastructure, and science.AWSThe AWS MCP Server is now generally availableAWS announces the general availability of the AWS MCP Server, a managed remote Model Context Protocol (MCP) server that gives AI agents and coding assistants secure, authenticated access to all AWS services. The AWS MCP…SlackFrom SSH to REST: A Security-Driven Modernization of Slack’s EMR Data PipelinesExcerpt By 2024, Slack s data platform had accumulated 700+ SSH-based operators orchestrating critical data pipelines. We re talking daily search indexing that processed terabytes of data, analytics jobs powering busines…MozillaTrustworthy JavaScript for the Open WebThe open web is a critical platform for applications that handle highly sensitive data, from private communications to financial transactions and medical records. Traditionally, servers are trusted to deliver the appropr…MicrosoftAzure Cosmos DB Conf 2026 Recap: Lessons from ProductionA team was running at 100% RU utilization. Throttles were compounding into retries. P99 latency was degrading. The assumption was obvious: provision more throughput. They didn’t. Instead, they found a single logical part…AWSModernize your workflows: Amazon WorkSpaces now gives AI agents their own desktop (preview)Amazon WorkSpaces now lets AI agents securely operate legacy desktop applications—without APIs or modernization—using IAM authentication, MCP support, and computer vision within existing security frameworks.AirbnbMonitoring reliably at scaleDesigning monitoring that works when everything else doesn’t. By : Abdurrahman J. Allawala Introduction When an incident hits, teams lean on observability to answer the only questions that matter: what’s broken, and why?…NetflixDemocratizing Machine Learning at Netflix: Building the Model Lifecycle GraphSaish Sali , Nipun Kumar , Sura Elamurugu Introduction As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When…NetflixDemocratizing Machine Learning at Netflix: Building the Model Lifecycle GraphSaish Sali , Nipun Kumar , Sura Elamurugu Introduction As Netflix has grown, machine learning continues to support our ability to deliver value to members and drive excellence across multiple areas of our business. When…InstacartEmpowering Carrot Ads with Domain Adaptive LearningAuthors: Trey Zhong, Xiyu Wang Contributors: Joseph Haraldson, Sharad Gupta, Sarah Lamacchia Introduction Carrot Ads is Instacart’s omnichannel retail media solution that allows retailer partners to build and scale their…InstacartEmpowering Carrot Ads with Domain Adaptive LearningAuthors: Trey Zhong, Xiyu Wang Contributors: Joseph Haraldson, Sharad Gupta, Sarah Lamacchia Introduction Carrot Ads is Instacart’s omnichannel retail media solution that allows retailer partners to build and scale their…DiscordDiscord Patch Notes: May 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.AWSAWS Weekly Roundup: What’s Next with AWS 2026, Amazon Quick, OpenAI partnership, and more (May 4, 2026)Last week, I took some time off in York, England, often described as the most haunted city in the country. I wandered through the ruins of abbeys that have stood for nearly a thousand years, walked along medieval walls,…SpotifyBuilding a Natural Language Interface to the Spotify Ads API with Claude Code PluginsTurning OpenAPI spec and Markdown files into a conversational ads management tool — no compiled code required. The post Building a Natural Language Interface to the Spotify Ads API with Claude Code Plugins appeared first…PinterestOptimizing ML Workload Network Efficiency (Part I): Feature TrimmerGuangtong Bai | Staff Software Engineer, Product ML Infrastructure*; Shantam Shorewala | Software Engineer II, Product ML Infrastructure*; Chi Zhang | Staff Software Engineer, AI Platform*; Neha Upadhyay | Software Engin…NetflixState of Routing in Model ServingBy Nipun Kumar , Rajat Shah , Peter Chng Introduction This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences…NetflixState of Routing in Model ServingBy Nipun Kumar , Rajat Shah , Peter Chng Introduction This is the first blog post in a multi-part series that shares technical insights into how our ML model serving infrastructure powers several personalized experiences…MetaHow Meta Is Strengthening End-to-End Encrypted BackupsThe HSM-based Backup Key Vault Meta s HSM-based Backup Key Vault provides the foundation for end-to-end encrypted backups for WhatsApp and Messenger. The system allows people to protect their backed-up message history wi…GrabData Mesh at Grab (Part II): The foundational tools behind certificationIntroduction In Part I , we discussed why Grab is investing in a data mesh, referred to as the Signals Marketplace within Grab, as part of our evolving data culture. We also explained how data certification aids teams in…Google DeepMindEnabling a new model for healthcare with AI co-clinicianResearching the path to AI-augmented care and development of an AI co-clinician.StripeEverything we announced at Sessions 2026We’re making Stripe even more programmable; protecting and propelling your business with the strength of the Stripe network; and building economic infrastructure for AI.StripeGiving agents the ability to payLink’s wallet for agents gives agents programmatic access to Link, including the ability to generate a one-time-use card or Shared Payment Token (SPT) backed by the cards and bank accounts already in your wallet. It’s bu…DiscordYou’ve Got (Too Much) Mail: Behind the Scenes of the 3/25/26 Voice OutageOn March 25th, voice and video on Discord suffered major degradation beginning at 12:13 PDT, lasting a little over three hours. Learn how the issue originated, how it affected systems across Discord, how we recovered, an…AWSTop announcements of the What’s Next with AWS, 2026At the "What's Next with AWS" 2026 event, AWS launched Amazon Quick—an AI assistant for work with a desktop app and expanded integrations—and expanded Amazon Connect into four agentic AI solutions for supply chain, hirin…AirbnbSkipper: Building Airbnb’s embedded workflow engineHow Airbnb built a lightweight workflow engine to solve durable execution. By : Ricardo Gamba , Andriy Sergiyenko Introduction: The durable execution problem Picture this hypothetical flow: A host submits an insurance cl…PinterestFrom Clicks to Conversions: Architecting Shopping Conversion Candidate Generation at PinterestAuthors: Richard Huang | Machine Learning Engineer II; Yu Liu | Senior Machine Learning Engineer; Ziwei Guo | Senior Machine Learning Engineer; Andy Mao | Staff Machine Learning Engineer; Supeng Ge | Sr. Staff Machine Le…Mistral AIWorkflows for work that runs the businessGoogle DeepMindAnnouncing our partnership with the Republic of KoreaGoogle DeepMind and Korea partner to accelerate scientific breakthroughs using frontier AI modelsAWSAWS Weekly Roundup: Anthropic & Meta partnership, AWS Lambda S3 Files, Amazon Bedrock AgentCore CLI, and more (April 27, 2026)Late March took me to Seattle for the Specialist Tech Conference, one of the most energizing gatherings of AWS specialists from around the world. It was an incredible opportunity to connect with peers, exchange experienc…NetflixScaling Camera File Processing at NetflixOrchestrating Media Workflows Through Strategic Collaboration Authors: Eric Reinecke , Bhanu Srikanth Introduction to Content Hub’s Media Production Suite At Netflix, we want to provide filmmakers with the tools they nee…NetflixScaling Camera File Processing at NetflixOrchestrating Media Workflows Through Strategic Collaboration Authors: Eric Reinecke , Bhanu Srikanth Introduction to Content Hub’s Media Production Suite At Netflix, we want to provide filmmakers with the tools they nee…DiscordMeasure Less to Learn More: Using Fewer, Higher-quality Metrics to Capture What MattersToo many experiment metrics can make meaningful changes harder to detect. Learn how Discord used simulations and Principal Component Analysis to maximize signal and reduce noise.CohereCohere and Aleph Alpha Join ForcesMicrosoftLangChain.js for Beginners: A Free Course to Build Agentic AI Apps with JavaScriptWant to build AI agents with JavaScript that go beyond basic chat completions? Agents that reason, call tools, and pull from knowledge bases on their own? We put together a free, open source course to help you get there.…LyftHow We Built a Smarter Pickup Experience for Gated CommunitiesIf you live in a gated community, you’ve been there: You request a ride from your apartment complex, expect your driver to come to you as usual, and then — your driver’s car icon just stops right at the front gate. You w…LyftHow We Built a Smarter Pickup Experience for Gated CommunitiesIf you live in a gated community, you’ve been there: You request a ride from your apartment complex, expect your driver to come to you as usual, and then — your driver’s car icon just stops right at the front gate. You w…YelpHow Yelp Keeps Server-Driven UI Consistent Across Four PlatformsIf you’ve read our earlier post, you already know about CHAOS—the server-driven UI (SDUI) framework we built at Yelp that powers our dynamic views. Until now, we’ve explored its architecture, backend implementation, and…SpotifyBackground Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4)How we used Honk, Backstage, and Fleet Management to ease the pain of migrating thousands of datasets. The post Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations (Honk, Part 4) appeared first…Google DeepMindDecoupled DiLoCo: A new frontier for resilient, distributed AI trainingMetaModernizing the Facebook Groups Search to Unlock the Power of Community KnowledgeWe’ve fundamentally transformed Facebook Groups Search to help people more reliably discover, sort through, and validate community content that’s most relevant to them. We’ve adopted a new hybrid retrieval architecture a…Google DeepMindPartnering with industry leaders to accelerate AI transformationGoogle DeepMind partners with global consultancies to bring the power of frontier AI to organizations around the world.CohereWhy MoE Models Get More From Speculative DecodingAirbnbBuilding a fault-tolerant metrics storage system at AirbnbHow we built a storage system that ingests 50 million samples per second and stores 2.5 petabytes of logical time series data. By : Rishabh Kumar Modern observability practice encourages instrumenting every meaningful co…PinterestSmarter URL Normalization at Scale: How MIQPS Powers Content Deduplication at PinterestShanhai Liao | Senior Software Engineer, Content Acquisition and Media Platform; Di Ruan, | Senior Staff Software Engineer, Content Acquisition and Media Platform; Evan Li, | Senior Engineering Manager, Content Acquisiti…BAIR (Berkeley)Gradient-based Planning for World Models at Longer Horizons.grasp-results-table table { font-size: 0.875rem; line-height: 1.35; width: 100%; } .grasp-results-table th, .grasp-results-table td { padding: 0.35rem 0.5rem; } /* Consistent whitespace between major sections (this post…NetflixThe Human Infrastructure: How Netflix Built the Operations Layer Behind Live at ScaleBy: Brett Axler , Casper Choffat , and Alo Lowry In the three years since our first Live show, Chris Rock: Selective Outrage , we have witnessed an incredible expansion of our live content slate and the live operations t…NetflixThe Human Infrastructure: How Netflix Built the Operations Layer Behind Live at ScaleBy: Brett Axler , Casper Choffat , and Alo Lowry In the three years since our first Live show, Chris Rock: Selective Outrage , we have witnessed an incredible expansion of our live content slate and the live operations t…MetaPost-Quantum Cryptography Migration at Meta: Framework, Lessons, and TakeawaysWe’re sharing lessons learned from Meta’s post-quantum cryptography (PQC) migration to help other organizations strengthen their resilience as industry transitions to post-quantum cryptography standards. We’re proposing…MetaCapacity Efficiency at Meta: How Unified AI Agents Optimize Performance at HyperscaleWe re sharing insights into Meta s Capacity Efficiency Program, where we ve built an AI agent platform that helps automate finding and fixing performance issues throughout our infrastructure. By leveraging encoded domain…DiscordMaking Discord on Desktop Look Just Right: Display Settings to Ease the EyesLearn all sorts of toggles, options, and features on Discord’s desktop app to help you view media at your pace, lower the strength of colors across the app, and make app content easier to see.PinterestFinding zombies in our systems: A real-world story of CPU bottlenecksVaibhav Shankar; Staff Software Engineer | Raymond Lee; Staff Software Engineer | Chia-Wei Chen; Staff Software Engineer | Shunyao Li; Sr. Software Engineer | Yi Li; Staff Software Engineer | Ambud Sharma; Principal Engi…Google DeepMindGemini 3.1 Flash TTS: the next generation of expressive AI speechOur newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.AirbnbPrivacy-first connections: Empowering social experiences at AirbnbDiscover how Airbnb prioritizes user privacy while building a more connected community, empowering guests to engage socially, connect confidently, and maintain control of their personal data. By: Joy Jing ✨ Building a mo…SlackManaging context in long-run agentic applicationsExcerpt In complex, long-running agentic systems, maintaining alignment and coherent reasoning between agents requires careful design. In this second article of our series, we explore these challenges and the mechanisms…PinterestScaling Recommendation Systems with Request-Level DeduplicationAuthors: Matt Lawhon | Sr. Machine Learning Engineer; Filip Ryzner | Machine Learning Engineer II; Kousik Rajesh | Machine Learning Engineer II; Chen Yang | Sr. Staff Machine Learning Engineer; Saurabh Vishwas Joshi | Pr…Google DeepMindGemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoningGemini Robotics ER 1.6: Enhancing spatial reasoning and multi-view understanding for autonomous robotics.NetflixEvaluating Netflix Show Synopses with LLM-as-a-Judgeby Gabriela Alessio , Cameron Taylor , and Cameron R. Wolfe Introduction When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — there are thousands of titles —…NetflixEvaluating Netflix Show Synopses with LLM-as-a-Judgeby Gabriela Alessio , Cameron Taylor , and Cameron R. Wolfe Introduction When members log into Netflix, one of the hardest choices is what to watch. The challenge isn’t a lack of options — there are thousands of titles —…PinterestPerformance for EveryoneAuthor: Lin Wang (Android Performance Engineer) Default Feature For mobile apps, performance is considered as the “default feature”, which means apps are expected to run fast and be responsive. It’s just as if we expect…YelpZero downtime Upgrade: Yelp’s Cassandra 4.x Upgrade StoryThe Database Reliability Engineering team at Yelp seamlessly upgraded more than a thousand Cassandra nodes with zero downtime. This post takes you behind the scenes of our upgrade strategy, from planning sessions to flaw…PinterestEvolution of Multi-Objective Optimization at Pinterest Home feedHomefeed: Jiacong He, Dafang He, Jie Cheng (former), Andreanne Lemay, Mostafa Keikha, Rahul Goutam, Dhruvil Deven Badani, Dylan Wang Content Quality: Jianing Sun, Qinglong Zeng ML Serving: Li Tang Introduction In feed re…AirbnbBuilding a high-volume metrics pipeline with OpenTelemetry and vmagentA production-tested approach for moving a large-scale metrics pipeline from StatsD to OpenTelemetry and Prometheus. By: Eugene Ma , Natasha Aleksandrova When migrating to a new monitoring system, you’ll want to frontload…NetflixStop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix ScaleBy Ben Sykes In a previous post , we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experienc…NetflixStop Answering the Same Question Twice: Interval-Aware Caching for Druid at Netflix ScaleBy Ben Sykes In a previous post , we described how Netflix uses Apache Druid to ingest millions of events per second and query trillions of rows, providing the real-time insights needed to ensure a high-quality experienc…DiscordDiscord Patch Notes: April 6, 2026Check out the finer details of the more technical fixes implemented into Discord recently.DropboxImproving storage efficiency in Magic Pocket, our immutable blob storeBy turning compaction into a layered, adaptive pipeline and strengthening our monitoring and controls, we made Magic Pocket more resilient to workload changes.Google DeepMindGemma 4: Byte for byte, the most capable open modelsGemma 4: Our most intelligent open models to date, purpose-built for advanced reasoning and agentic workflows.CohereThe Enterprise AI Maturity ModelDiscordMULTIPLAYER SEQUEL TO ACCLAIMED AAAA GAME “THE LAST MEADOW” ANNOUNCED: PLAYABLE NOWBand together with Discordians from across the world in Last Meadow Online, the world’s first DBMMIRPG. Available to play until April 7, 2026.SlackFrom Custom to Open: Scalable Network Probing and HTTP/3 Readiness with PrometheusThe Problem: Legacy Tooling and Its Limitations Currently, Slack utilizes a hybrid approach to network measurement, incorporating both internal (such as traffic between AWS Availability Zones) and external (monitoring tr…LyftPredicting Rider Conversion in Sparse Data Environments with Bayesian TreesAt Lyft, understanding how riders go through our user experience is fundamental to operating a healthy marketplace. Specifically, it is important to have a robust model determining if a rider will actually request a ride…LyftPredicting Rider Conversion in Sparse Data Environments with Bayesian TreesAt Lyft, understanding how riders go through our user experience is fundamental to operating a healthy marketplace. Specifically, it is important to have a robust model determining if a rider will actually request a ride…Google DeepMindReimagining the mouse pointer for the AI eraGoogle DeepMind is transforming the mouse pointer into a context-aware AI partner. Move beyond the friction of traditional prompting with intuitive AI collaboration in Chrome and beyond.YelpBuilding Biz Ask Anything: From Prototype to ProductIntroduction Users have access to a wealth of information on Yelp business pages – from reviews and photos to structured information, menus, and Ask the Community feature on the business page, a single business page can…DiscordHow Multi-Factor Authentication Helps Keep Your Discord Account SafeA Discord account is more than just your username and avatar. That’s why it’s important to help keep your account safe and secure by using Multi-Factor Authentication, SMS Backup Authentication & QR Code Login. Learn how…Google DeepMindGemini 3.1 Flash Live: Making audio AI more natural and reliableOur latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.MozillaFirefox Developer Edition and Beta: Try out Mozilla’s .rpm package!In January, we introduced our Nightly package for RPM-based Linux distributions. Today, we are thrilled to announce it is now available for Firefox Beta! Firefox Beta is great for testing your sites in a version of Firef…LyftBeyond A/B Testing: Using Surrogacy and Region-Splits to Measure Long-Term Effects in MarketplacesImage generated with Gemini 3 Pro (Google), 2026. Written by Amber Wang and Y oonji Kim at Lyft. Background Whenever you use the Lyft app, there is a complex balancing act happening behind the scenes. Various levers are…LyftBeyond A/B Testing: Using Surrogacy and Region-Splits to Measure Long-Term Effects in MarketplacesImage generated with Gemini 3 Pro (Google), 2026. Written by Amber Wang and Y oonji Kim at Lyft. Background Whenever you use the Lyft app, there is a complex balancing act happening behind the scenes. Various levers are…DropboxReducing our monorepo size to improve developer velocityMonorepos will continue to grow as products evolve, but growth doesn’t have to mean friction.Google DeepMindLyria 3 Pro: Create longer tracks in moreIntroducing Lyria 3 Pro, which unlocks longer tracks with structural awareness. We’re also bringing Lyria to more Google products and surfaces.Google DeepMindProtecting people from harmful manipulationGoogle DeepMind researches AI's harmful manipulation risks across areas like finance and health, leading to new safety measures.DiscordDiscord Update: March 24, 2026 ChangelogHere s the Discord Changelog from March 24, 2026, so you can stay informed on what’s new in recent app updates!Mistral AISpeaking of VoxtralEtsyMaking Ads Count: Using MMoE and Auxiliary Tasks to Better Connect Buyers & SellersWhen buyers search on Etsy, they need to quickly and easily find the perfect item. At the same time, sellers need to be confident their unique products are being seen by the right customers. Our Ads Search ranking model,…SlackHow Slack Rebuilt Notifications 📣Introduction 🔔 At Slack, notifications are how teams stay in the loop, but they can also become overwhelming when not designed with intention. Our goal was to make staying informed feel effortless. We set out to rebuild…GrabFrom firefighting to building: How AI agents restored our team’s core productivityAbstract Grab’s Analytics Data Warehouse (ADW) team supports over 1,000 users each month. These users support an extensive repository of more than 15,000 tables, which powers approximately 50% of all queries within our d…EtsyMigrating Etsy’s database sharding to VitessEtsy has maintained a sharded MySQL architecture since around 2010. This database cluster contains most of Etsy’s online data and is made up of ~1,000 tables distributed across ~1,000 shards. Over the last 16 years, it h…Mistral AIIntroducing ForgeDropboxHow we optimized Dash's relevance judge with DSPyWe used DSPy to turn prompt engineering for our relevance judge into a measurable, automated optimization loop, improving task performance, cost, and how reliably it works in production.Google DeepMindMeasuring progress toward AGI: A cognitive frameworkWe’re introducing a framework to measure progress toward AGI, and launching a Kaggle hackathon to build the relevant evaluations.Mistral AIMistral AI partners with NVIDIA to accelerate open frontier modelsMistral AILeanstral: Open-Source foundation for trustworthy vibe-codingMistral AIIntroducing Mistral Small 4DiscordHow ROOST is Advancing Online SafetyThe threat landscape online has shifted dramatically. Many online platforms are left to reinvent safety tools from scratch. That’s the gap ROOST was built to close — and it’s why open-sourcing battle-tested tools like Os…BAIR (Berkeley)Identifying Interactions at Scale for LLMs--> Understanding the behavior of complex machine learning systems, particularly Large Language Models (LLMs), is a critical challenge in modern artificial intelligence. Interpretability research aims to make the decisio…GrabEnabling R8 optimization at scale with AI-assisted debuggingGrab is Southeast Asia’s leading superapp, providing a suite of services that bring essential needs to users throughout the region. Its offerings include ride-hailing, food delivery, parcel delivery, mobile payments, and…DiscordYou’re Now Discord Official: Developers, Claim Your Game and Verify Your ServerDevelopers can now claim and customize their game’s profiles on Discord. Curate your game’s presence on the platform to help people discover more about your game, and get your server verified in the process! Read on to s…Mistral AIRails testing on autopilot: Building an agent that writes what developers won'tDiscordBuilding on the Social Layer of Games: What’s New from GDC 2026At GDC 2026, Discord gives developers more ways to close the gap between connection and play.GrabReclaiming Terabytes: Optimizing Android image caching with TLRUIntroduction In a previous post, we discussed Project Bonsai , our initiative to reduce the Grab app’s download size. We successfully reduced the Android Application Package (APK) download size by 26%. This reduction off…DiscordDiscord Patch Notes: March 6, 2026Check out the finer details of the more technical fixes implemented into Discord recently.DiscordTracing Discord's Elixir Systems (Without Melting Everything)Join Senior Software Engineer Nick Krichevsky as he explains how Discord added distributed tracing to Elixir s message passing and optimized it to handle millions of concurrent users.MozillaWhy is WebAssembly a second-class language on the web?This post is an expanded version of a presentation I gave at the 2025 WebAssembly CG meeting in Munich. WebAssembly has come a long way since its first release in 2017. The first version of WebAssembly was already a grea…InstacartOur Early Journey to Transform Instacart’s Discovery Recommendations with LLMsKey Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. O…InstacartOur Early Journey to Transform Instacart’s Discovery Recommendations with LLMsKey Contributors: Moein Hasani, Hamidreza Shahidi, Trace Levinson, Guanghua Shu Introduction At Instacart, we are laser-focused on improving the user experience by making shopping feel easy, engaging, and personalized. O…DropboxUsing LLMs to amplify human labeling and improve Dash search relevanceHow we train Dash's search ranking models with a mix of human and LLM-assisted labeling.MozillaGoodbye innerHTML, Hello setHTML: Stronger XSS Protection in Firefox 148Cross-site scripting (XSS) remains one of the most prevalent vulnerabilities on the web. The new standardized Sanitizer API provides a straightforward way for web developers to sanitize untrusted HTML before inserting it…DiscordGetting Global Age Assurance Right: What We Got Wrong and What's ChangingDiscord s CTO addresses community concerns about age assurance: no mass ID collection, new vendor transparency commitments, and a delayed global launch until second half of 2026.LyftScaling Localization with AI at LyftWritten by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scale…LyftScaling Localization with AI at LyftWritten by Stefan Zier For years, Lyft’s localization infrastructure relied exclusively on human translation. While this model usually ensured excellent quality, it was bound by multi-day turnarounds and costs that scale…DiscordHow to Change Your Theme to Bring Your Vibe to DiscordAdd a splash of personality and make Discord pop by changing Discord’s color scheme! Learn how to adjust the look of Discord on both desktop and mobile.DiscordOsprey: Open Sourcing our Rule EngineDiscord uses Osprey to quickly detect and remove new types of harm from putting our customers at risk. Now we’re open-sourcing this tool so others can do the same.InstacartTurning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with CapsightKey Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitiv…InstacartTurning Data into Velocity: Caper’s Edge and Cloud Data Flywheel with CapsightKey Contributors: Youming Luo, Andrew Tanner, Matas Sriubiskis, Sylvia Lin, Sikun Zhu, Lei Li, Xiao Zhou Introduction Caper is Instacart’s AI-powered smart cart that provides customers with a fast, seamless, and intuitiv…MozillaLaunching Interop 2026The Interop Project is a cross-browser initiative to improve web compatibility in areas that offer the most benefit to both users and developers. The group, including Apple, Google, Igalia, Microsoft, and Mozilla, takes…LyftTrusting the Untestable: Validation and Diagnostics for the Doubly Robust Modelswritten by Ross Chu and Shima Nassiri The Causal Frontier: Measurement Beyond Randomization The gold standard for determining the causal impact of a policy or product change at a company like Lyft is the A/B test (random…LyftTrusting the Untestable: Validation and Diagnostics for the Doubly Robust Modelswritten by Ross Chu and Shima Nassiri The Causal Frontier: Measurement Beyond Randomization The gold standard for determining the causal impact of a policy or product change at a company like Lyft is the A/B test (random…DropboxHow low-bit inference enables efficient AIMaking products like Dropbox Dash accessible to individuals and businesses means tackling new challenges around efficiency and resource use.DropboxInsights from our executive roundtable on AI and engineering productivityFrom Claude Code to Cursor, we're big adopters of AI coding tools at Dropbox. The early results have been promising, but there are still a lot of open questions about how to work with these tools most effectively and whe…InstacartFrom print to digital: Making weekly flyers shoppable at Instacart through computer vision and LLMsFrom Print to Digital: Making Weekly Flyers Shoppable at Instacart Through Computer Vision and LLMs Key contributors: Prithvi Srinivasan, Shishir Kumar Prasad, Kristen Morgan, Bryan Pham, Rick Shukla, Preeti Chadha, Vipu…InstacartFrom print to digital: Making weekly flyers shoppable at Instacart through computer vision and LLMsFrom Print to Digital: Making Weekly Flyers Shoppable at Instacart Through Computer Vision and LLMs Key contributors: Prithvi Srinivasan, Shishir Kumar Prasad, Kristen Morgan, Bryan Pham, Rick Shukla, Preeti Chadha, Vipu…Mistral AIVoxtral transcribes at the speed of sound.Mistral AIVoxtral transcribes at the speed of sound.DiscordDiscord Patch Notes: February 4, 2026Check out the finer details of the more technical fixes implemented into Discord recently.InstacartMigrating to Jetpack ComposeMigrating to Jetpack Compose: How AI Accelerated Our Journey at Caper Introduction At Instacart, our Caper smart carts bring together AI, computer vision, and real-time data to power the future of in-store shopping. Cust…InstacartMigrating to Jetpack ComposeMigrating to Jetpack Compose: How AI Accelerated Our Journey at Caper Introduction At Instacart, our Caper smart carts bring together AI, computer vision, and real-time data to power the future of in-store shopping. Cust…YelpHow Yelp Built a Back-Testing Engine for Safer, Smarter Ad Budget AllocationIntroduction Modern advertising platforms are fast-paced and interconnected: even small adjustments can have ripple effects on how ads are shown, how budgets are spent, and the value advertisers get from their ad spend.…DiscordHow to Customize Your Discord ProfileMake your first impression count with a profile that represents you how YOU want to be seen. Learn how to edit and customize your profile to have it rep you the right way.GrabCursor at Grab: Adoption and impactAdoption overview The illustration below encapsulates how Cursor is scaled across Grab, achieving rapid and widespread adoption that accelerated software development and empowered non-technical teams to build solutions.…DropboxEngineering VP Josh Clemm on how we use knowledge graphs, MCP, and DSPy in DashEngineering VP Josh Clemm deep-dives into how we think about knowledge graphs, indexes, MCP, and prompt optimization using tools like DSPy.Mistral AITerminally online Mistral Vibe.Mistral AITerminally online Mistral Vibe.Mistral AIHeaps do lie: debugging a memory leak in vLLM.EtsyHow Etsy Uses LLMs to Improve Search RelevanceEver searched for something specific, only to be met with results that are close, but not quite ? On Etsy’s Search Relevance team, that frustration is exactly what we are tackling. Our goal is simple yet ambitious: to he…BAIR (Berkeley)Information-Driven Design of Imaging SystemsAn encoder (optical system) maps objects to noiseless images, which noise corrupts into measurements. Our information estimator uses only these noisy measurements and a noise model to quantify how well measurements disti…LyftLyft’s Feature Store: Architecture, Optimization, and EvolutionWritten by Rohan Varshney , with support from Devon Mittow Janice Lee . This article expands upon a presentation from the Feature Store Summit 2025, which can be viewed in full here . There is also another video availabl…LyftLyft’s Feature Store: Architecture, Optimization, and EvolutionWritten by Rohan Varshney , with support from Devon Mittow Janice Lee . This article expands upon a presentation from the Feature Store Summit 2025, which can be viewed in full here . There is also another video availabl…Mistral AIIntroducing Mistral OCR 3LyftFrom Python3.8 to Python3.10: Our Journey Through a Memory LeakImage generated with ChatGPT (OpenAI), 2025. Intro When working with Python, memory management often feels like a solved problem. The garbage collector quietly does its job, and unlike C or C++, we rarely think about mal…LyftFrom Python3.8 to Python3.10: Our Journey Through a Memory LeakImage generated with ChatGPT (OpenAI), 2025. Intro When working with Python, memory management often feels like a solved problem. The garbage collector quietly does its job, and unlike C or C++, we rarely think about mal…DiscordHow to Make and Use Custom Emoji on DiscordEmojis on Discord are special — you can make a little picture out of almost any symbol, in-joke, or bizarre late-night inspiration.Mistral AIIntroducing: Devstral 2 and Mistral Vibe CLI.DiscordGift Ideas for the Dedicated Discord User in Your LifeWe’ve got plenty of gift ideas, from those who game through the night, to those who decorate their profile just right — let’s dig in!DiscordDiscord Patch Notes: December 8, 2025Check out the finer details of the more technical fixes implemented into Discord recently.DiscordYour Discord Checkpoint is Rolling Out! Celebrate What You Did in 2025How many messages did you send? How long did you hang out in voice? Who’d you talk with the most? Stop by your end-of-year Checkpoint, a recap of the stuff YOU did on Discord throughout the year!Mistral AIIntroducing Mistral 3DiscordBringing In-Game Commerce to Discord CommunitiesBy bringing commerce directly to official game communities, we re giving game developers the opportunity to benefit from these dynamics, creating incremental revenue that complements their existing storefronts.DiscordSave and Display Your Faves: Add Discord Shop & Marvel Rivals Items to Your Profile’s WishlistKeep track of all the stuff from the Shop you’ve been wanting to purchase with the new Wishlist feature. Display the stuff you’ve been eyeing on your profile, and if you’re lucky enough, maybe one of your friends may see…SlackStreamlining Security Investigations with AgentsSlack’s Security Engineering team is responsible for protecting Slack’s core infrastructure and services. Our security event ingestion pipeline handles billions of events per day from a diverse array of data sources. Rev…EtsyReducing experiment duration with predicted control variatesI n 2021, we published a blog post titled “ Increasing experimentation accuracy and speed by using control variates ,” describing how we reduce the variance of metrics using CUPED in our experimentation platform. This is…SlackAndroid VPAT journeyBackground A Voluntary Product Accessibility Template (VPAT) is a document that outlines how well a product aligns with accessibility (a11y) standards. Its primary purpose is to inform customers about a product s a11y fe…Mistral AIMistral AI - KI für DeutschlandLyftLyftLearn Evolution: Rethinking ML Platform ArchitectureWritten by Yaroslav Yatsiuk At Lyft, machine learning (ML) is the engine behind our most critical business functions — from dispatch and pricing optimization to fraud detection and support automation. Our ML infrastructu…LyftLyftLearn Evolution: Rethinking ML Platform ArchitectureWritten by Yaroslav Yatsiuk At Lyft, machine learning (ML) is the engine behind our most critical business functions — from dispatch and pricing optimization to fraud detection and support automation. Our ML infrastructu…DiscordHow to Share What You’re Playing, Listening to, or Watching as Your Status on DiscordPlaying a game right now? Listening to some tunes, or catching up on that one anime your friends won’t stop talking about? Learn how to show off what you’re up to as your Discord status and show @everyone what’s up!DiscordHow to Link Discord to Battlefield 6, Marvel Rivals & MoreSome of the most popular multiplayer games have added the ability to directly link your Discord account to the game! Learn how to link your Discord account to some big-name titles and see what sorta perks it provides.InstacartBuilding The Intent Engine: How Instacart is Revamping Query Understanding with LLMsAuthors: Yuanzheng Zhu, Guanghua Shu, Raochuan Fan, Vinesh Gudla, Tejaswi Tenneti Introduction When people search for items on Instacart, they don’t always type perfectly worded phrases. They might write “bread no gluten…InstacartBuilding The Intent Engine: How Instacart is Revamping Query Understanding with LLMsAuthors: Yuanzheng Zhu, Guanghua Shu, Raochuan Fan, Vinesh Gudla, Tejaswi Tenneti Introduction When people search for items on Instacart, they don’t always type perfectly worded phrases. They might write “bread no gluten…DiscordA Cornucopia of Updates Make Discord on Desktop Fresher Than a Crisp Fall BreezeThis fall, emoji making gets faster, the Settings page gets a redesign, Group DMs are easier to customize, more games gain support for special Discord-powered capabilities, and Family Center gets some expanded features.…DiscordDiscord Update: November 6, 2025 ChangelogHere s the Discord Changelog from November 6, 2025, so you can stay informed on what’s new in recent app updates!BAIR (Berkeley)RL without TD learningIn this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (…EtsyImproving performance by prefetching product pages from Etsy SearchRarely are there opportunities for big, bold, game-changing improvements in web performance. The Speculation Rules API (SRA) is a recent browser development that offers just such an opportunity. This post details a joint…Mistral AIIntroducing Mistral AI Studio.EtsyUnderstanding Etsy’s Vast Inventory with LLMsFor more than 20 years, Etsy has been the destination for human creativity online. Our marketplace is home to more than 100 million special items made, handpicked and designed by more than 5 million sellers. These items…EtsyUnlocking Faster Insights with Experimenter-Defined SegmentationsImagine you have a fabulous idea to drive more sales on Etsy by giving out free ice cream with every purchase. How would you know if it will actually work? One way to test this out is to run an experiment ! An experiment…YelpS3 server access logs at scaleIntroduction Yelp heavily relies on Amazon S3 (Simple Storage Service) to store a wide variety of data, from images, logs, database backups, and more. Since data is stored on the cloud, we need to carefully manage how th…Mistral AIMistral AI raises 1.7B€ to accelerate technological progress with AIEtsyBuilding Etsy Buyer Profiles with LLMsEvery day, shoppers from Etsy's community of nearly 90M buyers visit our marketplace to search for unique, handmade, and vintage items. But with over 100 million listings, how do we help each buyer find exactly what they…Mistral AILe Chat. Custom MCP connectors. Memories.Mistral AIMake Memory work for you.BAIR (Berkeley)What exactly does word2vec learn?What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task. Despite the fact that word2vec is a well-known prec…MozillaCRLite: Fast, private, and comprehensive certificate revocation checking in FirefoxFirefox is now the first and the only browser to deploy fast and comprehensive certificate revocation checking that does not reveal your browsing activity to anyone (not even to Mozilla). Tens of millions of TLS server c…EtsyContext engineering case studies: Etsy-specific question answeringThis post investigates the benefits and limitations of prompt engineering in two instances of AI-assisted onboarding relying on large language model (LLM) technology. Of particular interest is how truthful (and therefore…Mistral AIUnlocking the potential of vision language models on satellite imagery through fine-tuningMistral AIAnnouncing Codestral 25.08 and the Complete Mistral Coding Stack for EnterpriseMistral AIOur contribution to a global environmental standard for AIMistral AILe Chat dives deep.Mistral AIVoxtralMistral AIUpgrading agentic coding capabilities with the new Devstral modelsMistral AIUpgrading agentic coding capabilities with the new Devstral modelsYelpExploring CHAOS: Building a Backend for Server-Driven UIA little while ago, we published a blog post on CHAOS: Yelp’s Unified Framework for Server-Driven UI. We strongly recommend reading that post first to gain a solid understanding of SDUI and the goals of CHAOS. This post…Mistral AIAnnouncing AI for CitizensMistral AIAnnouncing AI for CitizensBAIR (Berkeley)Whole-Body Conditioned Egocentric Video Prediction.modal { display: none; position: fixed; z-index: 9999; padding-top: 50px; left: 0; top: 0; width: 100%; height: 100%; overflow: auto; background-color: rgba(0,0,0,0.9); } .modal-content { margin: auto; display: block; m…YelpRevenue Automation Series: Testing an Integration with Third-Party SystemBackground As described in the second blog post of Revenue Automation series, Revenue Data Pipeline processes a large amount of data via complex logic transformations to recognize revenue. Thus, developing a robust produ…PayPalPayPal Releases Agentic Toolkit to Accelerate CommerceThe following is a repost from the PayPal Developer Blog . Building on the release of PayPal’s MCP servers , PayPal is excited to introduce the PayPal Agentic Toolkit *. This toolkit empowers developers to seamlessly int…BAIR (Berkeley)Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs have improved, so have the attacks against them. Prompt injection attack is listed as the #1 threat by OWASP t…BAIR (Berkeley)Repurposing Protein Folding Models for Generation with Latent DiffusionPLAID is a multimodal generative model that simultaneously generates protein 1D sequence and 3D structure, by learning the latent space of protein folding models. The awarding of the 2024 Nobel Prize to AlphaFold2 marks…PayPalPayPal Begins Rollout of MCP Servers to Accelerate Agentic CommerceThe following is a repost from the PayPal Developer Blog by Prakhar Mehrotra, SVP of Artificial Intelligence, PayPal At PayPal, we strive to make it easier for developers to access our services. Today, we are taking the…EtsyBehind the Scenes - A Glimpse into Tax CalculationsIn the past, sellers were responsible for managing and fulfilling their own tax obligations. However, more and more jurisdictions are now requiring marketplaces such as Etsy to collect the tax from buyers and remit the t…MozillaImproving Firefox Stability in the Enterprise by Reducing DLL InjectionBeginning in version 138, Firefox will offer an alternative to DLL injection for Data Loss Prevention (DLP) deployments in enterprise environments. DLL Injection DLL injection into Firefox is a topic we’ve covered on the…BAIR (Berkeley)Scaling Up Reinforcement Learning for Traffic Smoothing: A 100-AV Highway DeploymentTraining Diffusion Models with Reinforcement Learning We deployed 100 reinforcement learning (RL)-controlled cars into rush-hour highway traffic to smooth congestion and reduce fuel consumption for everyone. Our goal is…PayPalEstimating Incremental Lift in Customer Value (Delta CV) using Synthetic ControlHow we measure the impact of user actions and product adoptions at PayPal In today’s competitive digital landscape, understanding user interactions with your products is essential for driving revenue and building lasting…MozillaLaunching Interop 2025Interop 2025 continues the mission to make the web more consistent across browsers, building on 2024’s 95% interoperability score. This year, 19 focus areas target key developer needs and long-standing issues, including…EtsyAdopting Jetpack Compose for Etsy’s Android AppOne of our Guiding Principles at Etsy is that we “commit to our craft.” This means that we have a culture of learning, in which we’re constantly looking for opportunities to improve and learn, adopt industry best practic…MozillaIntroducing Uniffi for React Native: Rust-Powered Turbo ModulesMozilla and Filament have introduced Uniffi for React Native, a tool that allows developers to leverage the safety and performance benefits of Rust in cross-platform React Native apps. The post Introducing Uniffi for Rea…MozillaLlamafile v0.8.14: a new UI, performance gains, and moreDiscover the latest release of Llamafile 0.8.14, an open-source AI tool by Mozilla Builders. With a new command-line chat interface, enhanced performance, and support for powerful models, Llamafile makes it easy to run l…Mozilla0Din: A GenAI Bug Bounty Program – Securing Tomorrow’s AI TogetherAs AI continues to evolve, so do the threats against it. As these GenAI systems become more sophisticated and widely adopted, ensuring their security and ethical use becomes paramount. 0Din is a groundbreaking GenAI bug…MozillaAnnouncing Official Puppeteer Support for FirefoxWe’re pleased to announce that, as of version 23, the Puppeteer browser automation library now has first-class support for Firefox. This means that it’s now easy to write automation and perform end-to-end testing using P…EtsyMachine Learning in Content Moderation at EtsyAt Etsy, we’re focused on elevating the best of our marketplace to help creative entrepreneurs grow their businesses. We continue to invest in making Etsy a safe and trusted place to shop, so sellers’ extraordinary items…MozillaSnapshots for IPC FuzzingProcess separation remains one of the most important parts of the Firefox security model and securing our IPC (Inter-Process Communication) interfaces is crucial to keep privileges in the different processes separated. W…MozillaSponsoring sqlite-vec to enable more powerful Local AI applicationsToday we’re proud to announce the next Mozilla Builders project: sqlite-vec. Led by independent developer Alex Garcia, this project brings vector search functionality to the beloved SQLite embedded database. Alex has bee…EtsyEnhancing Cloud Usage Forecasting, Monitoring & OptimizingIn 2020, Etsy concluded its migration from an on-premise data center to the Google Cloud Platform (GCP). During this transition, a dedicated team of program managers ensured the migration's success. Post-migration, this…EtsyEfficient Visual Representation Learning And EvaluationEtsy features a diverse marketplace of unique handmade and vintage items. It’s a visually diverse marketplace as well, and computer vision has become increasingly important to Etsy as a way of enhancing our users’ shoppi…MozillaExperimenting with local alt text generation in Firefox NightlyFirefox 130 will introduce an experimental new capability to automatically generate alt-text for images using a fully private on-device AI model. The feature will be available as part of Firefox’s built-in PDF editor, an…PayPalScaling PayPal’s AI Capabilities with PayPal Cosmos.AI PlatformBy Jun Yang , Zhenyin Yang , and Srinivasan Manoharan , based on the AI/ML modernization journey taken by the PayPal Cosmos.AI Platform team in the past three years. Source: Dall-E 3 AI is a transformative technology tha…MozillaLlamafile’s progress, four months inWhen Mozilla’s Innovation group first launched the llamafile project late last year, we were thrilled by the immediate positive response from open source AI developers. It’s become one of Mozilla’s top three most-favorit…MozillaPorting a cross-platform GUI application to RustIn this blog post, we delve into the motivations for choosing Rust for our crash reporter, outline the unique challenges of designing an application that operates when the main browser has failed, and discuss the new arc…MozillaPrototype even faster with the Gradio UI for Figma component libraryIn the fast-paced world of generative AI, staying ahead means moving swiftly and smartly. That's why we've embraced Gradio, the low-code prototyping toolkit from Hugging Face, as our go-to for bringing new ideas to life.…GoogleGenerative AI to quantify uncertainty in weather forecastingPosted by Lizao (Larry) Li, Software Engineer, and Rob Carver, Research Scientist, Google Research Accurate weather forecasts can have a direct impact on people’s lives, from helping make routine decisions, like what to…GoogleAutoBNN: Probabilistic time series forecasting with compositional bayesian neural networksPosted by Urs Köster, Software Engineer, Google Research Time series problems are ubiquitous, from forecasting weather and traffic patterns to understanding economic trends. Bayesian approaches start with an assumption a…GoogleUsing AI to expand global access to reliable flood forecastsPosted by Yossi Matias, VP Engineering Research, and Grey Nearing, Research Scientist, Google Research Floods are the most common natural disaster , and are responsible for roughly $50 billion in annual financial damages…GoogleComputer-aided diagnosis for lung cancer screeningPosted by Atilla Kiraly, Software Engineer, and Rory Pilgrim, Product Manager, Google Research Lung cancer is the leading cause of cancer-related deaths globally with 1.8 million deaths reported in 2020. Late diagnosis d…GoogleSCIN: A new resource for representative dermatology imagesPosted by Pooja Rao, Research Scientist, Google Research Health datasets play a crucial role in research and medical education, but it can be challenging to create a dataset that represents the real world. For example, d…GoogleScreenAI: A visual language model for UI and visually-situated language understandingPosted by Srinivas Sunkara and Gilles Baechler, Software Engineers, Google Research Screen user interfaces (UIs) and infographics, such as charts, diagrams and tables, play important roles in human communication and huma…GoogleMELON: Reconstructing 3D objects from images with unknown posesPosted by Mark Matthews, Senior Software Engineer, and Dmitry Lagun, Research Scientist, Google Research A person's prior experience and understanding of the world generally enables them to easily infer what an object lo…EtsyMacramé: Untangling the Knot on the Etsy Android Listing ScreenEasily the most important and complex screen in the Buy on Etsy Android app is the listing screen, where all key information about an item for sale in the Etsy marketplace is displayed to buyers. Far from just a title an…GoogleHEAL: A framework for health equity assessment of machine learning performancePosted by Mike Schaekermann, Research Scientist, Google Research, and Ivor Horn, Chief Health Equity Officer Director, Google Core Health equity is a major societal concern worldwide with disparities having many causes.…GoogleCappy: Outperforming and boosting large multi-task language models with a small scorerPosted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-…GoogleTalk like a graph: Encoding graphs for large language modelsPosted by Bahare Fatemi and Bryan Perozzi, Research Scientists, Google Research Imagine all the things around you — your friends, tools in your kitchen, or even the parts of your bike. They are all connected in different…GoogleChain-of-table: Evolving tables in the reasoning chain for table understandingPosted by Zilong Wang, Student Researcher, and Chen-Yu Lee, Research Scientist, Cloud AI Team People use tables every day to organize and interpret complex information in a structured, easily accessible format. Due to th…GoogleHealth-specific embedding tools for dermatology and pathologyPosted by Dave Steiner, Clinical Research Scientist, Google Health, and Rory Pilgrim, Product Manager, Google Research There’s a worldwide shortage of access to medical imaging expert interpretation across specialties in…GoogleSocial learning: Collaborative learning with large language modelsPosted by Amirkeivan Mohtashami, Research Intern, and Florian Hartmann, Software Engineer, Google Research Large language models (LLMs) have significantly improved the state of the art for solving tasks specified using n…GoogleCroissant: a metadata format for ML-ready datasetsPosted by Omar Benjelloun, Software Engineer, Google Research, and Peter Mattson, Software Engineer, Google Core ML and President, MLCommons Association Machine learning (ML) practitioners looking to reuse existing datas…EtsyHow We Built The Deals Tab in Swift UIBalancing Engineering Ambition with Product Realism Introduction In July of 2023, Etsy’s App Updates team, responsible for the Updates feed in Etsy’s mobile apps, set off with an ambitious goal: to revamp the Updates tab…GoogleGoogle at APS 2024Posted by Kate Weber and Shannon Leon, Google Research, Quantum AI Team Today the 2024 March Meeting of the American Physical Society (APS) kicks off in Minneapolis, MN. A premier conference on topics ranging across phys…GoogleVideoPrism: A foundational visual encoder for video understandingPosted by Long Zhao, Senior Research Scientist, and Ting Liu, Senior Staff Software Engineer, Google Research An astounding number of videos are available on the Web, covering a variety of content from everyday moments p…PayPalLeveraging Spark 3 and NVIDIA’s GPUs to Reduce Cloud Cost by up to 70% for Big Data PipelinesBy Ilay Chen and Tomer Akirav At PayPal, hundreds of thousands of Apache Spark jobs run on an hourly basis, processing petabytes of data and requiring a high volume of resources. To handle the growth of machine learning…GoogleAdvances in private training for production on-device language modelsPosted by Zheng Xu, Research Scientist, and Yanxiang Zhang, Software Engineer, Google Language models (LMs) trained to predict the next word given input text are the key technology for many applications [ 1 , 2 ]. In Gbo…GoogleLearning the importance of training data under concept driftPosted by Nishant Jain, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research The constantly changing nature of the world around us poses a significant challenge for the development of AI model…GoogleDP-Auditorium: A flexible library for auditing differential privacyPosted by Mónica Ribero Díaz, Research Scientist, Google Research Differential privacy (DP) is a property of randomized mechanisms that limit the influence of any individual user’s information while processing and analyz…GoogleGraph neural networks in TensorFlowPosted by Dustin Zelle, Software Engineer, Google Research, and Arno Eigenwillig, Software Engineer, CoreML Objects and their relationships are ubiquitous in the world around us, and relationships can be as important to…GoogleIntervening on early readouts for mitigating spurious features and simplicity biasPosted by Rishabh Tiwari, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research Machine learning models in the real world are often trained on limited data that may contain unintended statistic…GoogleA decoder-only foundation model for time-series forecastingPosted by Rajat Sen and Yichen Zhou, Google Research Time-series forecasting is ubiquitous in various domains, such as retail, finance, manufacturing, healthcare and natural sciences. In retail use cases, for example, it…GoogleMobileDiffusion: Rapid text-to-image generation on-devicePosted by Yang Zhao, Senior Software Engineer, and Tingbo Hou, Senior Staff Software Engineer, Core ML Text-to-image diffusion models have shown exceptional capabilities in generating high-quality images from text prompt…GoogleMixed-input matrix multiplication performance optimizationsPosted by Manish Gupta, Staff Software Engineer, Google Research AI-driven technologies are weaving themselves into the fabric of our daily routines, with the potential to enhance our access to knowledge and boost our ov…GoogleExphormer: Scaling transformers for graph-structured dataPosted by Ameya Velingker, Research Scientist, Google Research, and Balaji Venkatachalam, Software Engineer, Google Graphs , in which objects and their relations are represented as nodes (or vertices) and edges (or links…PayPalDeclarative Feature Engineering at PayPalPhoto by fabio on Unsplash PayPal supports over 400 million active consumers and merchants worldwide. Every minute there are several thousand payment transactions. To prevent fraud in real-time at such a scale, we need t…PayPalStreamlining Developer Productivity with the PayPal Visual Studio Code ExtensionIn the ever-evolving landscape of software development, productivity and efficiency have become paramount to success. Developers are constantly juggling multiple tasks, from navigating complex codebases to integrating th…PayPalManaging Recurring Payments with Apple Pay Using PayPalRecurring payments have become an integral part of the modern digital economy, offering convenience and predictability for both consumers and businesses. Our previous post highlighted different methods of integrating App…PayPalAccept E-Commerce Payments Easily with PayPal’s Buttons ComponentAccepting online payments is now a universal must-have, catering to everyone from solo entrepreneurs to massive global corporations. PayPal’s Standard Checkout allows for seamless integration of PayPal’s Payment Buttons…PayPalWhy You Should Attend PayPal’s Developer Meetup at Money20/20The world of technology is constantly evolving, and developers are at the forefront of this dynamic landscape. Staying updated on the latest trends, tools, and innovations is not just a choice but a necessity for those i…EtsyThe AR Measuring Box: Etsy's answer to Big Tape MeasureA little while ago, Etsy introduced a new feature in its iOS app that could place Etsy sellers' artwork on a user's wall using Apple's Augmented Reality (AR) tools. It let them visualize how a piece would look in their s…EtsyThe So-fine Real-time ML ParadigmIntroduction Each year, Etsy hosts an event known as “CodeMosaic” - an internal hackathon in which Etsy admin propose and build bold advances quickly in our technology across a number of different themes. People across E…EtsyLeveraging Real-Time User Actions to Personalize Etsy AdsIntroduction Personalization is vital to connect our unique marketplace to the right buyer at the right time. Etsy has recently introduced a novel, general approach to personalizing ML models based on encoding and learni…Stanford AI LabLinkBERT: Improving Language Model Training with Document LinkLanguage Model Pretraining Language models (LMs), like BERT 1 and the GPT series 2 , achieve remarkable performance on many natural language processing (NLP) tasks. They are now the foundation of today’s NLP systems. 3 T…Stanford AI LabStanford AI Lab Papers and Talks at ACL 2022The 60th Annual Meeting of the Association for Computational Linguistics (ACL) 2022 is taking place May 22nd - May 27th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to pape…Stanford AI LabStanford AI Lab Papers and Talks at ICLR 2022The International Conference on Learning Representations (ICLR) 2022 is being hosted virtually from April 25th - April 29th. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to…Stanford AI LabDiscovering the systematic errors made by machine learning modelsDiscovering systematic errors with cross-modal embeddings In this blog post, we introduce Domino, a new approach for discovering systematic errors made by machine learning models. We also discuss a framework for quantita…Stanford AI LabGrading Complex Interactive Coding Programs with Reinforcement Learning[Summary] tl;dr: A tremendous amount of effort has been poured into training AI algorithms to competitively play games that computers have traditionally had trouble with, such as the retro games published by Atari, Go, D…Stanford AI LabUnderstanding Deep Learning Algorithms that Leverage Unlabeled Data, Part 1: Self-trainingDeep models require a lot of training examples, but labeled data is difficult to obtain. This motivates an important line of research on leveraging unlabeled data, which is often more readily available. For example, larg…Stanford AI LabStanford AI Lab Papers and Talks at AAAI 2022The 36th AAAI Conference on Artificial Intelligence (AAAI 2022) is being hosted virtually from February 22th - March 1st. We’re excited to share all the work from SAIL that’s being presented, and you’ll find links to pap…Stanford AI LabHow to Improve User Experience (and Behavior): Three Papers from Stanford's Alexa Prize TeamIntroduction In 2019, Stanford entered the Alexa Prize Socialbot Grand Challenge 3 for the first time, with its bot Chirpy Cardinal , which went on to win 2nd place in the competition. In our previous post , we discussed…Stanford AI LabReward Isn't Free: Supervising Robot Learning with Language and Video from the WebThis work was conducted as part of SAIL and CRFM . Deep learning has enabled improvements in the capabilities of robots on a range of problems such as grasping 1 and locomotion 2 in recent years. However, building the qu…Stanford AI LabBanditPAM: Almost Linear-Time k-medoids Clustering via Multi-Armed BanditsTL;DR Want something better than \(k\)-means? Our state-of-the-art \(k\)-medoids algorithm from NeurIPS, BanditPAM, is now publicly available! \(\texttt{pip install banditpam}\) and you're good to go! Like the \(k\)-mean…Stanford AI LabStanford AI Lab Papers and Talks at NeurIPS 2021The thirty-fifth Conference on Neural Information Processing Systems (NeurIPS) 2021 is being hosted virtually from Dec 6th - 14th. We’re excited to share all the work from SAIL that’s being presented at the main conferen…Stanford AI LabStanford AI Lab Papers at CoRL 2021The Conference on Robot Learning (CoRL 2021) will take place next week. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs below. Feel free to reach ou…Stanford AI LabStanford AI Lab Papers at EMNLP/CoNLL 2021The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021) will take place next week, colocated with CoNLL 2021. We’re excited to share all the work from SAIL that will be presented, and you’ll…Stanford AI LabSelective Classification Can Magnify Disparities Across GroupsSelective classification, where models are allowed to “abstain” when they are uncertain about a prediction, is a useful approach for deploying models in settings where errors are costly. For example, in medicine, model e…Stanford AI LabStanford AI Lab Papers at ICCV 2021The International Conference on Computer Vision (ICCV 2021) will be hosted virtually next week. We’re excited to share all the work from SAIL that will be presented, and you’ll find links to papers, videos and blogs belo…