Across frontier models, gpt-5.3-codex achieves the best overall performance (solving 19/22 tasks, 86.4%), outperforming claude-opus-4.6 (15/22, 68.2%), and kimi-2.5 exhibits the strongest performance among open-source models
The firm is also whitelisting a handful of market makers, including longtime crypto liquidity provider Wintermute, to facilitate trading. Meanwhile, access to BUIDL is restricted to qualified purchasers, a legal designation for those with assets of $5 million or more.
For years, Trump has claimed he had “no idea” about Epstein’s abuse of underage girls. Yet records show that in 2006, he privately told Palm Beach police that “everyone” knew about Epstein’s activities and described Ghislaine Maxwell as evil.
Trump’s call to Palm Beach police chief
According to an FBI interview conducted in October 2019 with former Palm Beach Police Chief Michael Reiter, Trump personally called him in July 2006, just as Epstein’s criminal sex charges became public. Reiter told agents that Trump said, “Thank goodness you’re stopping him, everyone has known he’s been doing this.”
Observational Memory achieves the highest score ever recorded on LongMemEval — 94.87% with gpt-5-mini — while maintaining a completely stable, cacheable context window. It beats the oracle, outperforms complex multi-step reranking systems with a single pass, and scales better with model quality than existing approaches.
"I mean, there's tons of redacted stuff. ... And [Trump's] name, I think I put his name, and it appears more than a million times. So it's all over the place."
The bottom line: "To me, this whole rollout of saying that members can come from nine to five to sit at those four computers, is just part of the coverup," Raskin asserted.
The 3 million documents that the administration has not publicly released "are the ones I'd like to see," he said.
"The administration says that these are duplicative. Well go ahead and release them then! If they're duplicative, what's the problem? We'll be the judge of that." "Epstein's lawyers synopsized and quoted Trump as saying that Jeffrey Epstein was not a member of his club at Mar-a-Lago, but he was a guest at Mar-a-Lago, and he had never been asked to leave," Raskin said. "That was redacted for some indeterminate, inscrutable reason."
Among participants who use AI, we find a stark divide in skill formation outcomes between high scoring interaction patterns (65%-86% quiz score) vs low-scoring interaction patterns (24%-39% quiz score). The high scorers only asked AI conceptual questions instead of code generation or asked for explanations to accompany generated code; these usage patterns demonstrate a high level of cognitive engagement.
We develop a model of political cycles driven by time-varying risk aversion. Agents choose to work in the public or private sector and to vote Democratic or Republican. In equilibrium, when risk aversion is high, agents elect Democrats—the party promising more redistribution. The model predicts higher average stock market returns under Democratic presidencies, explaining the well-known “presidential puzzle.” The model can also explain why economic growth has been faster under Democratic presidencies. In the data, Democratic voters are more risk averse, and risk aversion declines during Democratic presidencies. Public workers vote Democratic, while entrepreneurs vote Republican, as the model predicts.
We may be on the descending portion of a productivity J-curve. As Brynjolfsson, Rock, and Syverson illustrate, when firms adopt transformative general-purpose technologies, measured productivity often initially falls because resources are diverted to investment, reorganization, and learning that do not show up as measured output.
The task-completion time horizon is the task duration (measured by human expert completion time) at which an AI agent is predicted to succeed with a given level of reliability
it will automatically set all users’ accounts to a “teen-appropriate” experience unless they demonstrate that they’re adults
Among those to leave OpenAI in recent months over the strategic shift are vice-president of research Jerry Tworek, model policy researcher Andrea Vallone and economist Tom Cunningham.
MaxRL is a framework that turns more compute into increasingly better approximations of the maximum likelihood objective in sampling-based tasks.
If this perspective is accurate then it has deep implications for the economics of AI: the marginal cost of solving an idiosyncratic problem is small (you just need to map it to one of the canonical problems, and apply that solution), but there’s very high value in making progress on the canonical problems. So we would expect AI labs to be spending huge amounts of compute on advancing the SoTA on the few deep problems of the world, and providing a service that solves idiosyncratic problems very cheaply.
Fuzzers work by throwing massive amounts of random inputs at code to see what breaks. Opus 4.6 reads and reasons about code the way a human researcher would—looking at past fixes to find similar bugs that weren't addressed, spotting patterns that tend to cause problems, or understanding a piece of logic well enough to know exactly what input would break it.
You wouldn't think that chess players, who sit for hours on end and extend their arms only from time to time, would struggle with weight loss. But they do. Inside the very real, very bizarre metabolic phenomenon gripping chess.
Although bias against men was shown by both men and women, another finding of the study was the bias was larger from women than from men in three out of four of their research questions.
Despite our culture voting for parties that enforce gender equality schemes that put men’s education and careers behind women’s, women very often want a man who earns more money and has a higher education than she does. Thus we are living in a situation that seems designed to leave everyone dissatisfied, yet this is the design that people demand.
Counterintuitively, that’s my biggest reason to be optimistic about AI and creativity. When hard parts become easy, the differentiator becomes love.
This systematic review and meta-analytic investigation found that SFV use was associated with poorer cognition (attention, inhibitory control, language, memory, and working memory) and most mental health indices except body image and self-esteem.