Coding & Terminal Tasks
Terminal-Bench 3 & 4
Terminal-oriented tasks for evaluating coding and tool-using agents in command-line workflows and software engineering environments.
- Samples
- Samples on request
- Spec
- shared on request

Data infrastructure for AI agents
Task-oriented datasets for coding agents, interactive environments, frontier reasoning, safety, and multilingual evaluation.
01tasks/fix-failing-build/task.toml
schema_version = "1.4"[metadata]category = "software-engineering"[agent]timeout_sec = 900.0[environment]network_mode = "no-network"
02Task files
03harbor run
04Verifier result
01Dataset catalogue
Task-oriented datasets for training, evaluating and benchmarking AI systems, from single terminal tasks to work that runs for hours.
10 of 10 datasets
02Solutions
Start from the evaluation problem, then narrow to the data. Each area links to the catalogue entries that serve it.
03Methodology
A reference structure for task-oriented evaluation data, readable by researchers and procurement teams alike. Use it to judge whether any dataset gives you what you need to measure an agent.
Stage 01 of 05
Define the objective, starting conditions, constraints, and expected outcome.
These stages describe a robust evaluation-data workflow in general terms. How a specific RCC dataset maps to each of them is something to confirm with the team as part of a sample request.
04Why RCC
A dataset is only useful if it measures what you care about. The catalogue and the request process are built around that question.
The catalogue is organised by what you need to evaluate: coding, interactive environments, frontier reasoning, safety and multilingual work.
Ask for samples of exactly the datasets you are considering, one at a time or several in a single request.
Each entry says what kind of task it covers, and separates what is documented from what is confirmed with you directly.
Talk through environment, tooling and evaluation requirements with the team before you commit to anything.
From catalogue entry to sample to a judgement about fit, without a long qualification survey in between.
schema_version = "1.4"artifacts = ["/app"] [task]name = "example/fix-failing-build" [metadata]category = "software-engineering"difficulty = "hard"tags = ["terminal", "python"] [agent]timeout_sec = 900.0 [verifier]timeout_sec = 300.0environment_mode = "separate" [environment]network_mode = "no-network"cpus = 2memory_mb = 409605About
RCC Data Services helps research and engineering teams working on coding agents, reinforcement learning environments, long-horizon tasks, benchmark evaluation, frontier reasoning, agent safety and multilingual capability source task-oriented data, and assess a sample before they commit.
Contact
The quickest route to the team is a sample request, or a short note about the requirement you're working on.

06Sample requests
Tell us which dataset you're exploring. We'll use your request to understand what you're looking for and follow up about sample availability.
What happens next
Four fields at most. No phone number, budget or long questionnaire.