
Summary
- On October 6, OpenAI published a case study on how Jump Trading uses ChatGPT and GPT-6 Astra to hand long, ambiguous quantitative research to agents.
- Lucas Baker, Head of LLM R&D at Jump Trading, said agents can now run for days, weave together many data sources and stack their gains through recursive improvement.
- Agent-generated trading signals are scoped and reviewed like any other signal, with humans handling final validation to manage risk in a regulated industry.
North American quantitative trading firm Jump Trading is handing market research that runs for days at a time to agents built on OpenAI's GPT-6 Astra. OpenAI released a customer case study on October 6 saying Jump Trading is using ChatGPT and GPT-6 Astra to take on longer, more ambiguous research problems. Lucas Baker, who leads LLM research and development at the firm, said that "with the GPT-6 series, especially GPT-6 Astra, OpenAI has unlocked a new tier of autonomy for long-horizon tasks." The point, he explained, is that work which once required frequent human intervention has shifted to defining a secure, well-monitored environment with clear goals.
Jump Trading's business is predicting asset prices even slightly better. The firm builds predictive models using market data, news, events and a range of alternative data. Markets are complex, noisy and change over time, which makes it hard to anticipate exactly what will happen. According to Baker, predicting even slightly better than a coin flip at scale is enough for a successful strategy. The firm trades across every time horizon and every asset class.
The scale shows in the data. According to an abstract Baker presented with colleague Loren Puchalla Fiore at an ICLR 2026 expo session on April 25 this year, Jump Trading processes terabytes of market data every day, searching for predictive signals across thousands of instruments worldwide. Its pipeline runs from raw order book events through NLP signal integration and model training to live execution, all under tight latency constraints. At that session the firm presented results from fine-tuning large language models on text inputs to produce live trading signals, and from multi-agent systems that combine unstructured sources into forecasts used by human traders and automated strategies alike.
How AI is used has changed sharply over the past year. According to OpenAI's case study, a tool that once wrote code snippets or found small bugs now builds entire codebases and services on its own. Baker's team finds AI works best when treated like a colleague. Once researchers define a key problem, a work environment and a way to evaluate the quality and significance of results, they steer one or many agents in real time on where to focus the analysis and which job to run next.
What changed with GPT-6 Astra is the ability to stack improvements. According to Baker, agents no longer stop at finding meaningful changes; they merge and layer those wins in a process of recursive improvement. Within a single long-running task, the system analyzes its own findings, judges them against criteria agreed in the initial proposal, and changes direction without a person reviewing every round. "You can define something that needs to run for days, that needs to pull from many data sources, that needs to make those subtle calls about what is important and what is not, and that needs to interrelate everything," Baker said, adding that "a comprehensive analysis like that actually works now."
Finance is a heavily regulated industry, so mistakes carry both financial and compliance costs. Jump Trading manages that risk with system design and boundaries, clear constraints, infrastructure that supports steering and observation, and human review of changes. "If you have a safe environment where the agent is free to produce any output it needs, but there is also a human review process at the end where critical validation takes place with human acceptance, that's what gives us confidence," Baker said. Even when an agent produces a trading signal, it is scoped and reviewed like any other output. It is treated as a single signal that is usually informative but potentially wrong, and combined with every other signal in a stringently reviewed and controlled execution environment. The ICLR abstract also covered hallucination risk management, strict point-in-time correctness and the evaluation methods used to ensure reliability in production.
From an AI engineer's point of view, the heart of this case is the harness more than the model. Baker's job itself is building the agents, harnesses and infrastructure that let researchers explore ideas more broadly and deeply. The human role shifts toward writing goals, evaluation metrics and environment definitions, while agents take on execution and iteration. Even so, GPT-6 Astra's longest-running work still involves regular check-ins with the person who defined the task. Humans confirm not only what data to pull, how long to run and what matters, but also whether intermediate results make sense. METAL has previously reported that OpenAI published a GPT-6 Astra case study with tax firm Basis, and also covered the GPT-6 model guide.
The next step Baker envisions is autoresearch. The idea is that recursive improvement of measurable systems by agent researchers will become part of a quantitative researcher's everyday workflow. Once humans set the inputs, environment, evaluation metrics and priorities, a loosely structured fleet of agents coordinated by other agents would handle the rest, deciding how to explore ideas, allocate compute time and integrate promising findings. In the original OpenAI case study, which METAL reviewed, Baker noted that in 2024 an early agent writing a single file without mistakes was impressive, that in 2025 building whole codebases from scratch became possible, and that in 2026 many agents working together made progress on open research questions. "If you can solve a Millennium problem, you can probably also figure out some pretty interesting facts about quant finance," he said.
The change this case shows is in the unit of research. Not a single question and answer, but days-long experiments are handed to agents, while people write the fences and the scorecard for those experiments. Jump Trading's review system, which treats any single signal as an input that may be wrong, is the floor that supports that shift. The question is moving from how long an agent can run on its own to how precisely human-defined criteria can hold that long run in place.





Comments