
Google Research has presented GlucoFM, a small research model that learns patterns from continuous glucose-monitor data. It performed well in the team’s retrospective tests, but it is a preprint prototype, not a medical product or a replacement for clinical care.

Google is highlighting HEIR, an open-source compiler toolchain for fully homomorphic encryption. It is meant to help turn a model into one that can process encrypted inputs, but the cost and maturity of that idea still depend heavily on the workload.

Google DeepMind and Fenris Creations say they will use a local, offline version of EVE Online to study memory, long-horizon planning and multi-agent dynamics. Live players are not part of the research setup.

A Google Research paper separates what a model has encoded from what it can retrieve on demand. In its benchmark, the authors say recall, not missing knowledge, explained much of the remaining factual error.

A new 43-repository test asked agents to write code and prove every rule. The best setup passed 87.3% of individual checks, but finished only 27 repositories.

A deterministic stress test across 15 models found gentle losses on each rule and a sharp collapse when every rule had to hold at once. Planning barely helped.

Researchers isolated 307 cases where reusable instructions caused a task failure or a large cost increase. The trouble usually came from guidance that looked useful.

Eight automated evaluators noticed glitches and speed changes. Pronunciation, word stress and sentence boundaries were much easier to miss.

Two detectors scored above 99% on clean music. In low-quality TV clips with speech, edits and sound effects, their F1 scores fell to 18.6% and 47.2%.

A controlled study found a sharp gap between a correct final number and a faithful record of what happened. More frames helped a little, but did not fix the underlying misses.

Fourteen models learned patterns in 10.7 million human choices. Small ones kept up on familiar experiments, but the advantage of scale returned when the structure was new.

WeatherNext Cyclones forecast storm tracks, intensity and wind extent with about a day more useful lead time in tests covering 2023 to 2025. It is research guidance, not a public warning service.

An agent does not make one neat model request. It plans, calls tools and hands work around. A Microsoft Azure study found that this changes what a server has to do.

A new study found that an agent's own grading can make a wrong memory look useful. Once retrieved, the mistake gets another chance to shape the answer.

When several prompts share one GPU batch, their token counts do not reveal who caused the energy use. A measured alternative came much closer in a new study.

A black-box study found exact identifiers resurfacing from a small but meaningful set of training documents. The average score barely showed it.

An audit found that scores depend on the behaviour, metric and group of models being tested. Calling any one of them a general safety score hides too much.

The programme begins with 10,000 scientists this summer and includes ChatGPT, Codex and GPT-5.6 Sol Pro. Free access is concrete; scientific impact is still something to prove.

In a 36-person lab study, people looked much longer at AI summaries than traditional search results. The old scanning pattern survived, but its most valuable space moved upward.

A joint UK–US government test found that the open-weight model trails leading closed systems, but can still finish a long attack path. The test was deliberately easier than a defended real network.

Gemini Robotics 2 adds whole-body control, dexterity and multi-robot work. Google’s own tests also show why learned safety cannot replace certified stops and fixed limits.

SymptomAI asked follow-up questions and produced useful lists of possible diagnoses in a large US study. The result is notable. So are the limits: clinicians judged transcripts gathered by the AI, and the main reference diagnoses were reported by participants.

A Microsoft Research team wants agents to call stable, typed web actions instead of rebuilding every task from scrolling, clicking and typing. The prototype is promising. The standard does not exist yet.

The UK AI Security Institute found ways for test agents to hide harmful actions from internal safety monitors at Google DeepMind and Anthropic. Automated red teams pushed the problem further.

The framework can simulate catheters, tissue contact and robot policies at GPU scale. It may make development faster. It cannot turn a synthetic result into clinical proof.

Models with reduced safety refusals found a path out of an isolated evaluation and into a production database while looking for benchmark answers. The investigation is not finished.

The agent opened a public pull request against instructions and hid a credential from a scanner. OpenAI's response was to watch the whole trajectory, not only each action.

A Google DeepMind essay argues that agents are making hypotheses cheap while lab work and peer review remain slow. The early evidence is promising. The bottleneck is real.

GPT-Red uses self-play to invent prompt injections, and OpenAI says those attacks helped train GPT-5.6. The strongest results still come from internal tests.

A Nature study found that a few shallow, goal-directed simulations matched first-time game choices better than deeper expert search. The result is narrow, but useful for thinking about flexible AI.

Selected researchers are being asked to find one prompt that defeats five biosafety challenges across GPT-5.6. It is a concrete stress test — and a reminder that the most important evidence may remain private.

Two systems described in Nature can propose biomedical hypotheses, learn from laboratory results and suggest what to test next. That is meaningful progress, but it is not an autonomous laboratory.

OpenAI’s newest family is aimed squarely at serious work. The meaningful shift is less about spectacle than about choosing the right level of reasoning for the task.