Can Jev Nav?
Have System One “LLMs” finally broken the real-time robotics frontier?
Loading task…
- Score
- —
- SPL
- —
- Run time
- —
- Driven
- —
Jev learns navigation.
The recent launch of Typesafe’s Jev have left many re-evaluating the ability for language models to play a role in real-time control. Are “System One” models running at 2-5 hz now able to apply its internet-scale pretraining to robotics tasks? Can they compete with algorithmic, deterministic approaches to path and motion planning? Or is the language-only nature of Jev and similar models an impossible barrier.
We put this to the test across both real and simulated environments, giving models short, medium, and long navigation tasks in a variety of environments on a quadruped robot.
The test / model matrix is a follows:
| Driver | Model | Harness | Tools | Robot access |
|---|---|---|---|---|
| Dimensional | Generic | Dimcode harness with navigation & path planning tool call access | Navigation and path planning tool calls | Prompted to be the control |
| Typesafe | Jev | TypeSafeAgent, one document in, six typed answers out, at 2 Hz | None: choices only | WorldState in, velocity out |
| Astra | gpt-6-astra | Pi | Bash, grep, plus Pi's native read, edit, write, find and ls; a shell in a container with Python, uv and any public library | Three Zenoh topics only, world_state, cmd_vel, finished, described in a README |
| Fable | claude-fable-5-1 | Pi | same as above | same as above |
| Opus | claude-opus-4-7 | Pi | same as above | same as above |
| GPT 5.6 | gpt-5.6-luna | Pi | same as above | same as above |
Test matrix.
| Home size | Floor | Homes | Tasks | Objects per home (mean) | Rooms (median) |
|---|---|---|---|---|---|
| Small (< 77 m²) | cramped | 32 | 44 | 240 | 2 |
| Small (< 77 m²) | open | 12 | 16 | 165 | 2 |
| Medium (77–207 m²) | cramped | 24 | 52 | 396 | 4 |
| Medium (77–207 m²) | open | 20 | 49 | 257 | 4 |
| Large (> 207 m²) | cramped | 11 | 31 | 488 | 5 |
| Large (> 207 m²) | open | 34 | 135 | 391 | 6 |
| All | 133 | 327 | 323 | 4 |
Scene selection
Scenes are 134 homes from the Habitat Dataset [1], run in Habitat-Sim [2][3], with doors removed so every room is reachable; homes were classified by navigable floor area into small, medium and large categories, under 77 m², 77 to 207 m², and over 207 m². Then categorized by how crowded the floor is, so we have an even distribution between cramped and open environments.
Task Selection
Tasks are all navigation objectives from start point to a target object. Objects are pulled from the scene graphs. In each environment, rooms are defined by flood-filling the navmesh. To be an eligible goal point, every object must have a navigable standing point within 1 m of its footprint and a viable navmesh route from the start point. One goal per room, up to six per home, qualified by sitting at least 3 m apart, adding at least 5 m of new route to the task, and overlapping no other task by more than 60%.
Perception.
Agents were provided with environments in JSON/string format, as Jev can only receive language as input. We define this abstraction as WorldState. We use the below method to convert from XML scenegraphs to WorldState, which is a string representation of the environment, objects, and obstacles.
From a home to a world state.
Jev needs language
Jev performs poorly when WorldState provides numerical values relative to world frame. For example:
40% reached34 of 84 tasks · mean score 0.40
"robot": {"position": {"x": 1.2, "y": -0.4}, "yaw_deg": 92},
"objects": [{"label": "chair", "position": {"x": 3.1, "y": 0.8},
"distance_m": 2.3, "bearing_deg": 31}, ... 20 objects]90% reached76 of 84 tasks · mean score 0.84 · 42 gained, 0 lost
"objects": [{"target": true, "label": "chair", "bearing": "ahead_left",
"distance": "near", "bearing_deg": 31}, ... 5 nearest],
"way_to_target": {"state": "blocked", "blocked_by": "wall",
"open_sides": [{"side": "right", "kind": "doorway", ...}]},
"robot": {"recent": {"pattern": "advancing"}}The above two experiments were run on a limited environment and task set of 11 scenes and 84 total tasks. The only change was in the input WorldState format as described: world state → robot state with natural language helpers
We found task completion rates increased from 40% to 90% when instead Jev received STRING values in robot frame, and with helper worlds in natural language such as ahead_left, behind_right, near, far, touching, blocked, tight, clear, doorway, corner, stuck and moving_without_getting_closer as well as bearings in units of degrees out of 360.
Memory.
Jev is stateless, meaning each call is independent unless otherwise prompted. The literature has yet to evaluate how Jev’s implicit memory degrades over time (its context limit is 32K tok for state and 64K tok including questions). With that low ceiling, it remains to be seen if memory can be injected effectively over long periods. However, for each model, WorldState gives the following values that cache short-term memory:
| Field | What it holds |
|---|---|
going_around | The side the robot chose to pass a blocker, with how long it has held it (for_s) |
robot.recent | An 8 s window of poses reduced to moved_m, turned_deg, target_closer_m and a pattern: starting, advancing, still, stuck, turning on the spot, or moving without getting closer |
been_there | The robot’s own trail, kept at 0.5 m spacing; any open side whose far point lies within 0.75 m of the trail is flagged been_there |
free_space | Implicit memory: it saves free space, so it implicitly logs space that has already been visited |
Trajectory Scoring.
We score trajectories by first defining a geodesic path G for every task which is the shortest navigable path from starting point to target object

F the navigable free space of the home, s the start, g the standing point beside the target; ℓ is the length of the shortest path inside F.Then we use a variant of SPL [4] which is the success weighted by path length and delta from the geodesic G. This is calculated per run for each model.

S arrival, 1 if the run ended within 1 m of the target with line of sight, else 0; ℓ the geodesic length; p the length the robot actually drove. A run that arrives along the geodesic scores 1; every metre of detour lowers it.This penalizes paths that diverge too far from the ideal path. Since S with SPL is binary success, S ∈ {0, 1}, we also grade with a SoftSPL [5] in which S is continuous as a function of final arrival distance from target object. Worth noting that current trajectory grading does not include time, smoothness, and other parameters.

d_final the navmesh geodesic from where the robot stopped to the goal, so a run that stops halfway keeps half the credit; an arrived run has d_final = 0 and the same value as SPL. (·)₊ clips at zero.Grading
In addition to trajectory scoring, we grade across the following:
| Metric | Note |
|---|---|
| # Collisions | counted when a drive comand of at least 0.1 m/s moves the robot less than 20% of what it asked for over 0.5 s, i.e. it is pushing against something |
| Time to target | Time elapsed during navigation to target object |
| Cost | input + output tokens |
| Path smoothness | trajectory scoring does not penalize jerky motion if that motion gets the robot closer to the target, so we include this to offset that fact |
| Arrived | yes/no |
| Excess Turning amount in radians/m | excess relative to the geodesic |
Runtime.
All models get a fresh dimOS instance and Habitat sim per task. The dimOS instance includes a Perception module DemoObjects which publishes a stream of Detection3DArray as a stand-in for real perception input from the robot. This is to normalize inputs for all models in the test matrix.
RawRobotBridge JSON (as described above) is then built from the dimOS instance with robot pose (PoseStamped), detections (Detection3DArray), 2D detections (Detection2DArray), pointcloud (PointCloud2), goal (str), the robot's own motion block (RobotState), and optionally a Memory object (Memory).
Jev
TypeSafeAgent calls WorldState at 2hz and sends output to Jev. Jev outputs the following at frequency f defined as 1 / max(0.5 s, l_j) where l_j is latency per jev call inference call. In plain language, if Jev answers in under 0.5 s the loop runs at 2 Hz. It may run slower than 2hz depending on inference time but speed is capped at 2hz.
Jev raw output
| Answer | Type | Choices |
|---|---|---|
drive.x | choice | forward, backward, or none |
drive.y | choice | left, right, or none |
drive.yaw | choice | turn_left, turn_right, or none |
stop | yes / no with a confidence | “should the robot stop right now” |
task | choice | finished or continue |
target | choice | which entry of objects the goal names, or none |
Output then converted to Drive object with signed values for x, y, and yaw, converted to dimOS Twist vector and then sent to the robot. As an aside there is additional logic we apply in the Twist conversion that helps smooth the application of velocity, but not worth elaborating on in depth here.
All other Agents
All agents with harnesses are not permitted to use dimOS tools, enforced by their only access to dimOS instance being RawRobotBridge. RawRobotBridge calls RawRobotBridge at 2hz and all agents access the current world state via the robot/world_state/json Zenoh published topic. All agents successfully wrote their own simple client, typically a Python script over eclipse-zenoh.
All agents for this run use Pi, except for Dimensional which runs dimcode. Tools provided Pi's native coding tools only: read, bash, edit, write, plus grep, find, ls.
RawRobotBridge exposes velocity commands, which agent could publish to, clamped to 1 m/s and 1.5 rad/s.
Jev on a Real Robot.
When brought into the real-world, multi-room navigation was a bottleneck. We believe that this is a data problem rather than a true limitation imposed by Jev, as the WorldState relied on clean labeling of walls and doorways which is more challenging outside of a simulated environment with ground truth.
What is clear is that there is opportunity for a control loop sitting comfortably at around 4hz - either by providing coarse inputs or by signalling for help to the slower but more intelligent LLM outer loops.
Results.
1. Astra wins, but Jev and Fable are close behind
Results at a glance
| Driver | Mean SPL | Arrived | SPL | SoftSPL | Time / ideal | Path / geodesic | Cost / run | Recorded |
|---|---|---|---|---|---|---|---|---|
01DimensionalDimcode with navigation skills | 0.743 | 88.4%289 / 327 | 0.743 | 0.774 | 1.40× | 1.22× | Negligible # calls* | 327 / 327 |
02Astracoding agent | 0.522 | 76.1%249 / 327 | 0.522 | 0.578 | 7.72× | 1.68× | $1.610 | 327 / 327 |
03Fable 5.1coding agent | 0.330 | 50.8%166 / 327 | 0.330 | 0.405 | 10.71× | 1.91× | $0.963 | 327 / 327 |
04TypeSafeJev, hosted | 0.263 | 45.9%150 / 327 | 0.263 | 0.325 | 4.24× | 2.55× | $0.081 | 327 / 327 |
05GPT-5.6coding agent | 0.213 | 32.1%105 / 327 | 0.213 | 0.330 | 10.21× | 1.78× | $0.094 | 327 / 327 |
06Opus 4.7coding agent | 0.125 | 18.7%61 / 327 | 0.125 | 0.262 | 8.28× | 2.43× | $0.378 | 327 / 327 |
*Dimcode agent will typically run a single Navigate() tool call
Arrival rate (%)
Mean SPL
Mean SoftSPL
2. Jev achieves Fable-level performance on shorter paths, and wins speed and cost
| Driver | Arrived | Median time when arrived | Cost per run |
|---|---|---|---|
| Dimensional | 84.4% 38 / 45 | 22 s | Negligible # calls* |
| TypeSafe | 71.1% 32 / 45 | 51 s | $0.055 |
| Astra | 88.9% 40 / 45 | 132 s | $0.877 |
| Fable 5.1 | 68.9% 31 / 45 | 143 s | $0.685 |
| GPT-5.6 | 57.8% 26 / 45 | 172 s | $0.057 |
| Opus 4.7 | 35.6% 16 / 45 | 81 s | $0.318 |
| Driver | Arrived | Median time when arrived | Time / ideal (median) | Cost per run |
|---|---|---|---|---|
| Dimensional | 88.4% | 50 s | 1.28× | Negligible # calls* |
| TypeSafe | 45.9% | 101 s | 3.05× | $0.081 |
| Astra | 76.1% | 228 s | 6.71× | $1.610 |
| Fable 5.1 | 50.8% | 249 s | 7.46× | $0.963 |
| GPT-5.6 | 32.1% | 262 s | 8.75× | $0.094 |
| Opus 4.7 | 18.7% | 123 s | 4.37× | $0.378 |
*Dimcode agent will typically run a single Navigate() tool call
3. Astra has the greatest consistency across path lengths & complexity
Reproduction.
# git clone https://github.com/dimensionalOS/dimos.git && cd dimos && git checkout feat/typesafe-world-state dimos evals run dimos.evals.suites.habitat_nav --agent dimos.evals.agents.topic --set 'modules=["type-safe-agent"]' --set trace=TypeSafeAgent --case 102343992_chair
# git clone https://github.com/dimensionalOS/dimos.git && cd dimos && git checkout feat/typesafe-world-state dimos evals run dimos.evals.suites.habitat_nav --agent dimos.evals.agents.pi --set no_dimos=true --set model=gpt-6-astra --case 102343992_chair
Cite this work.
If you use Can Jev Nav? in your research, please cite this report.
Stash Pomichter, Ruthwik Dasyam, and Henry Ventura. Can Jev Nav? Dimensional Research. https://research.dimensional.org/system-one-navigation/
@misc{dimensionalcanjevnav,
author = {Pomichter, Stash and Dasyam, Ruthwik and Ventura, Henry},
title = {{Can Jev Nav?}},
howpublished = {Online research report, Dimensional Research},
year = {2026},
url = {https://research.dimensional.org/system-one-navigation/}
}References.
- Khanna, M., Mao, Y., Jiang, H., Haresh, S., Schacklett, B., Batra, D., Clegg, A., Undersander, E., Chang, A. X., Savva, M. "Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation." CVPR 2024. https://arxiv.org/abs/2306.11290
- Savva, M. et al. "Habitat: A Platform for Embodied AI Research." ICCV 2019. https://arxiv.org/abs/1904.01201
- Szot, A. et al. "Habitat 2.0: Training Home Assistants to Rearrange their Habitat." NeurIPS 2021. https://arxiv.org/abs/2106.14405
- Anderson, P. et al. "On Evaluation of Embodied Navigation Agents." 2018. https://arxiv.org/abs/1807.06757
- Datta, S. et al. "Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents." CoRL 2020 (SoftSPL, Habitat Challenge). https://arxiv.org/abs/2009.03231