Dimensional Research / Navigation

Can Jev Nav?

Stash Pomichter, Ruthwik Dasyam, Henry Ventura

Have System One “LLMs” finally broken the real-time robotics frontier?

Explore sample runs
Scene 103997919

Loading task…

Loading recorded geometry and paths…
TypeSafeLoadingDashed: geodesic · coloured: recorded path
Score
—
SPL
—
Run time
—
Driven
—
327navigation tasks
133environments
6models with harnesses
1,962recorded runs
01

Jev learns navigation.

The recent launch of Typesafe’s Jev have left many re-evaluating the ability for language models to play a role in real-time control. Are “System One” models running at 2-5 hz now able to apply its internet-scale pretraining to robotics tasks? Can they compete with algorithmic, deterministic approaches to path and motion planning? Or is the language-only nature of Jev and similar models an impossible barrier.

We put this to the test across both real and simulated environments, giving models short, medium, and long navigation tasks in a variety of environments on a quadruped robot.

The test / model matrix is a follows:

DriverModelHarnessToolsRobot access
DimensionalGenericDimcode harness with navigation & path planning tool call accessNavigation and path planning tool callsPrompted to be the control
TypesafeJevTypeSafeAgent, one document in, six typed answers out, at 2 HzNone: choices onlyWorldState in, velocity out
Astragpt-6-astraPiBash, grep, plus Pi's native read, edit, write, find and ls; a shell in a container with Python, uv and any public libraryThree Zenoh topics only, world_state, cmd_vel, finished, described in a README
Fableclaude-fable-5-1Pisame as abovesame as above
Opusclaude-opus-4-7Pisame as abovesame as above
GPT 5.6gpt-5.6-lunaPisame as abovesame as above
02

Test matrix.

Home sizeFloorHomesTasksObjects per home (mean)Rooms (median)
Small (< 77 m²)cramped32442402
Small (< 77 m²)open12161652
Medium (77–207 m²)cramped24523964
Medium (77–207 m²)open20492574
Large (> 207 m²)cramped11314885
Large (> 207 m²)open341353916
All1333273234

Scene selection

Scenes are 134 homes from the Habitat Dataset [1], run in Habitat-Sim [2][3], with doors removed so every room is reachable; homes were classified by navigable floor area into small, medium and large categories, under 77 m², 77 to 207 m², and over 207 m². Then categorized by how crowded the floor is, so we have an even distribution between cramped and open environments.

Task Selection

Tasks are all navigation objectives from start point to a target object. Objects are pulled from the scene graphs. In each environment, rooms are defined by flood-filling the navmesh. To be an eligible goal point, every object must have a navigable standing point within 1 m of its footprint and a viable navmesh route from the start point. One goal per room, up to six per home, qualified by sitting at least 3 m apart, adding at least 5 m of new route to the task, and overlapping no other task by more than 60%.

03

Perception.

Agents were provided with environments in JSON/string format, as Jev can only receive language as input. We define this abstraction as WorldState. We use the below method to convert from XML scenegraphs to WorldState, which is a string representation of the environment, objects, and obstacles.

Inside one tick / Figure 2

From a home to a world state.

Recorded input · tick 12
chair · 11.5 mTarget bearing +2° · 11.5 mBlocked at 9.2 mwall · 9.2 mcorner · 65° detourcorner · 3.9 m cleardoorway · 11° detourdoorway · 1.8 m wideahead: 4.2 m · clear4.2 mahead_left: 3.5 m · clear3.5 mleft: 0.8 m · tightbehind_left: 0.7 m · tightbehind: 3 m · clear3.0 mbehind_right: 5 m · clear5.0 mright: 5 m · clear5.0 mahead_right: 5 m · clear5.0 mPool ladder: behind_left, 0.69 m, width 0.7 mTuuci Ocean Master Recta: behind_right, 2.91 m, width 3.7 mstool: behind_left, 3.06 m, width 0.6 mstool: behind_left, 3.27 m, width 0.6 mcurtain: ahead_left, 3.58 m, width 2.3 mrobot2 mHSSD / 105515286
Target distance11.5 m
Direct routeBlocked
Doorway width1.8 m
0.0 / 9s · 2×

Jev needs language

Jev performs poorly when WorldState provides numerical values relative to world frame. For example:

Experiment #1 · world frame

40% reached34 of 84 tasks · mean score 0.40

"robot": {"position": {"x": 1.2, "y": -0.4}, "yaw_deg": 92},
"objects": [{"label": "chair", "position": {"x": 3.1, "y": 0.8},
             "distance_m": 2.3, "bearing_deg": 31}, ... 20 objects]
Experiment #2 · robot frame, words

90% reached76 of 84 tasks · mean score 0.84 · 42 gained, 0 lost

"objects": [{"target": true, "label": "chair", "bearing": "ahead_left",
             "distance": "near", "bearing_deg": 31}, ... 5 nearest],
"way_to_target": {"state": "blocked", "blocked_by": "wall",
             "open_sides": [{"side": "right", "kind": "doorway", ...}]},
"robot": {"recent": {"pattern": "advancing"}}

The above two experiments were run on a limited environment and task set of 11 scenes and 84 total tasks. The only change was in the input WorldState format as described: world state → robot state with natural language helpers

We found task completion rates increased from 40% to 90% when instead Jev received STRING values in robot frame, and with helper worlds in natural language such as ahead_left, behind_right, near, far, touching, blocked, tight, clear, doorway, corner, stuck and moving_without_getting_closer as well as bearings in units of degrees out of 360.

04

Memory.

Jev is stateless, meaning each call is independent unless otherwise prompted. The literature has yet to evaluate how Jev’s implicit memory degrades over time (its context limit is 32K tok for state and 64K tok including questions). With that low ceiling, it remains to be seen if memory can be injected effectively over long periods. However, for each model, WorldState gives the following values that cache short-term memory:

FieldWhat it holds
going_aroundThe side the robot chose to pass a blocker, with how long it has held it (for_s)
robot.recentAn 8 s window of poses reduced to moved_m, turned_deg, target_closer_m and a pattern: starting, advancing, still, stuck, turning on the spot, or moving without getting closer
been_thereThe robot’s own trail, kept at 0.5 m spacing; any open side whose far point lies within 0.75 m of the trail is flagged been_there
free_spaceImplicit memory: it saves free space, so it implicitly logs space that has already been visited
05

Trajectory Scoring.

We score trajectories by first defining a geodesic path G for every task which is the shortest navigable path from starting point to target object

ℓ(s, g) = inf over paths γ in F with γ(0)=s and γ(1)=g of the integral from 0 to 1 of the norm of γ'(t) dtF the navigable free space of the home, s the start, g the standing point beside the target; ℓ is the length of the shortest path inside F.

Then we use a variant of SPL [4] which is the success weighted by path length and delta from the geodesic G. This is calculated per run for each model.

SPL = S times ℓ over max(p, ℓ)S arrival, 1 if the run ended within 1 m of the target with line of sight, else 0; ℓ the geodesic length; p the length the robot actually drove. A run that arrives along the geodesic scores 1; every metre of detour lowers it.

This penalizes paths that diverge too far from the ideal path. Since S with SPL is binary success, S ∈ {0, 1}, we also grade with a SoftSPL [5] in which S is continuous as a function of final arrival distance from target object. Worth noting that current trajectory grading does not include time, smoothness, and other parameters.

SoftSPL = max(0, 1 − d_final / ℓ) times ℓ over max(p, ℓ)d_final the navmesh geodesic from where the robot stopped to the goal, so a run that stops halfway keeps half the credit; an arrived run has d_final = 0 and the same value as SPL. (·)₊ clips at zero.

Grading

In addition to trajectory scoring, we grade across the following:

MetricNote
# Collisionscounted when a drive comand of at least 0.1 m/s moves the robot less than 20% of what it asked for over 0.5 s, i.e. it is pushing against something
Time to targetTime elapsed during navigation to target object
Costinput + output tokens
Path smoothnesstrajectory scoring does not penalize jerky motion if that motion gets the robot closer to the target, so we include this to offset that fact
Arrivedyes/no
Excess Turning amount in radians/mexcess relative to the geodesic
06

Runtime.

All models get a fresh dimOS instance and Habitat sim per task. The dimOS instance includes a Perception module DemoObjects which publishes a stream of Detection3DArray as a stand-in for real perception input from the robot. This is to normalize inputs for all models in the test matrix.

RawRobotBridge JSON (as described above) is then built from the dimOS instance with robot pose (PoseStamped), detections (Detection3DArray), 2D detections (Detection2DArray), pointcloud (PointCloud2), goal (str), the robot's own motion block (RobotState), and optionally a Memory object (Memory).

Jev

TypeSafeAgent calls WorldState at 2hz and sends output to Jev. Jev outputs the following at frequency f defined as 1 / max(0.5 s, l_j) where l_j is latency per jev call inference call. In plain language, if Jev answers in under 0.5 s the loop runs at 2 Hz. It may run slower than 2hz depending on inference time but speed is capped at 2hz.

Jev raw output

AnswerTypeChoices
drive.xchoiceforward, backward, or none
drive.ychoiceleft, right, or none
drive.yawchoiceturn_left, turn_right, or none
stopyes / no with a confidence“should the robot stop right now”
taskchoicefinished or continue
targetchoicewhich entry of objects the goal names, or none

Output then converted to Drive object with signed values for x, y, and yaw, converted to dimOS Twist vector and then sent to the robot. As an aside there is additional logic we apply in the Twist conversion that helps smooth the application of velocity, but not worth elaborating on in depth here.

All other Agents

All agents with harnesses are not permitted to use dimOS tools, enforced by their only access to dimOS instance being RawRobotBridge. RawRobotBridge calls RawRobotBridge at 2hz and all agents access the current world state via the robot/world_state/json Zenoh published topic. All agents successfully wrote their own simple client, typically a Python script over eclipse-zenoh.

All agents for this run use Pi, except for Dimensional which runs dimcode. Tools provided Pi's native coding tools only: read, bash, edit, write, plus grep, find, ls.

RawRobotBridge exposes velocity commands, which agent could publish to, clamped to 1 m/s and 1.5 rad/s.

07

Jev on a Real Robot.

When brought into the real-world, multi-room navigation was a bottleneck. We believe that this is a data problem rather than a true limitation imposed by Jev, as the WorldState relied on clean labeling of walls and doorways which is more challenging outside of a simulated environment with ground truth.

What is clear is that there is opportunity for a control loop sitting comfortably at around 4hz - either by providing coarse inputs or by signalling for help to the slower but more intelligent LLM outer loops.

08

Results.

1. Astra wins, but Jev and Fable are close behind

Six drivers · one task set

Results at a glance

DriverMean SPLArrivedSPLSoftSPLTime / idealPath / geodesicCost / runRecorded
01DimensionalDimcode with navigation skills
0.74388.4%289 / 3270.7430.7741.40×1.22×Negligible # calls*327 / 327
02Astracoding agent
0.52276.1%249 / 3270.5220.5787.72×1.68×$1.610327 / 327
03Fable 5.1coding agent
0.33050.8%166 / 3270.3300.40510.71×1.91×$0.963327 / 327
04TypeSafeJev, hosted
0.26345.9%150 / 3270.2630.3254.24×2.55×$0.081327 / 327
05GPT-5.6coding agent
0.21332.1%105 / 3270.2130.33010.21×1.78×$0.094327 / 327
06Opus 4.7coding agent
0.12518.7%61 / 3270.1250.2628.28×2.43×$0.378327 / 327

*Dimcode agent will typically run a single Navigate() tool call

Arrival rate (%)

TypeSafe: 45.9%, $0.08 per run. Astra: 76.1%, $1.61 per run. Fable 5.1: 50.8%, $0.96 per run. GPT-5.6: 32.1%, $0.09 per run. Opus 4.7: 18.7%, $0.38 per run.0$0.0025$0.4850$0.9775$1.45100$1.93Dimensional dimcode 88.4%TypeSafeAstraFable 5.1GPT-5.6Opus 4.7Mean cost / run (USD)
All recorded runs, including failures. Select a model for exact values.

Mean SPL

TypeSafe: 0.263, $0.08 per run. Astra: 0.522, $1.61 per run. Fable 5.1: 0.330, $0.96 per run. GPT-5.6: 0.213, $0.09 per run. Opus 4.7: 0.125, $0.38 per run.0.00$0.000.25$0.480.50$0.970.75$1.451.00$1.93Dimensional dimcode 0.743TypeSafeAstraFable 5.1GPT-5.6Opus 4.7Mean cost / run (USD)
All recorded runs, including failures. Select a model for exact values.

Mean SoftSPL

TypeSafe: 0.325, $0.08 per run. Astra: 0.578, $1.61 per run. Fable 5.1: 0.405, $0.96 per run. GPT-5.6: 0.330, $0.09 per run. Opus 4.7: 0.262, $0.38 per run.0.00$0.000.25$0.480.50$0.970.75$1.451.00$1.93Dimensional dimcode 0.774TypeSafeAstraFable 5.1GPT-5.6Opus 4.7Mean cost / run (USD)
All recorded runs, including failures. Select a model for exact values.
DimensionalNegligible # calls* · dashed reference88.4% arrived · 0.743 SPL

2. Jev achieves Fable-level performance on shorter paths, and wins speed and cost

Routes under 10 m · 45 tasks
DriverArrivedMedian time when arrivedCost per run
Dimensional84.4% 38 / 4522 sNegligible # calls*
TypeSafe71.1% 32 / 4551 s$0.055
Astra88.9% 40 / 45132 s$0.877
Fable 5.168.9% 31 / 45143 s$0.685
GPT-5.657.8% 26 / 45172 s$0.057
Opus 4.735.6% 16 / 4581 s$0.318
All 327 tasks · speed and cost
DriverArrivedMedian time when arrivedTime / ideal (median)Cost per run
Dimensional88.4%50 s1.28×Negligible # calls*
TypeSafe45.9%101 s3.05×$0.081
Astra76.1%228 s6.71×$1.610
Fable 5.150.8%249 s7.46×$0.963
GPT-5.632.1%262 s8.75×$0.094
Opus 4.718.7%123 s4.37×$0.378

*Dimcode agent will typically run a single Navigate() tool call

0%25%50%75%100%0×1×2×4×8×12×16×Tasks arrivedElapsed time / ideal time · ideal = geodesic at 0.5 m/s
Dimensional 88.1% 288/327TypeSafe 30.3% 99/327Astra 14.1% 46/327Fable 5.1 9.2% 30/327GPT-5.6 2.8% 9/327Opus 4.7 8.9% 29/327
Loading measured model calls…

3. Astra has the greatest consistency across path lengths & complexity

0%25%50%75%100%Tasks arrived<10 m45 tasksDimensional, <10 m: 84.4% arrived84TypeSafe, <10 m: 71.1% arrived71Astra, <10 m: 88.9% arrived89Fable 5.1, <10 m: 68.9% arrived69GPT-5.6, <10 m: 57.8% arrived58Opus 4.7, <10 m: 35.6% arrived3610–20 m129 tasksDimensional, 10–20 m: 89.1% arrived89TypeSafe, 10–20 m: 57.4% arrived57Astra, 10–20 m: 87.6% arrived88Fable 5.1, 10–20 m: 62.8% arrived63GPT-5.6, 10–20 m: 42.6% arrived43Opus 4.7, 10–20 m: 22.5% arrived2320–30 m69 tasksDimensional, 20–30 m: 84.1% arrived84TypeSafe, 20–30 m: 34.8% arrived35Astra, 20–30 m: 69.6% arrived70Fable 5.1, 20–30 m: 46.4% arrived46GPT-5.6, 20–30 m: 23.2% arrived23Opus 4.7, 20–30 m: 10.1% arrived1030 m+84 tasksDimensional, 30 m+: 92.9% arrived93TypeSafe, 30 m+: 23.8% arrived24Astra, 30 m+: 57.1% arrived57Fable 5.1, 30 m+: 26.2% arrived26GPT-5.6, 30 m+: 9.5% arrived10Opus 4.7, 30 m+: 10.7% arrived11
Hover, tap or focus a bar to inspect.
09

Reproduction.

Source on GitHubdimensionalOS / dimos
TYPESAFE
# git clone https://github.com/dimensionalOS/dimos.git && cd dimos && git checkout feat/typesafe-world-state
dimos evals run dimos.evals.suites.habitat_nav --agent dimos.evals.agents.topic --set 'modules=["type-safe-agent"]' --set trace=TypeSafeAgent --case 102343992_chair
CODING AGENT
# git clone https://github.com/dimensionalOS/dimos.git && cd dimos && git checkout feat/typesafe-world-state
dimos evals run dimos.evals.suites.habitat_nav --agent dimos.evals.agents.pi --set no_dimos=true --set model=gpt-6-astra --case 102343992_chair
10

Cite this work.

If you use Can Jev Nav? in your research, please cite this report.

Stash Pomichter, Ruthwik Dasyam, and Henry Ventura. Can Jev Nav? Dimensional Research. https://research.dimensional.org/system-one-navigation/

BIBTEX
@misc{dimensionalcanjevnav,
  author       = {Pomichter, Stash and Dasyam, Ruthwik and Ventura, Henry},
  title        = {{Can Jev Nav?}},
  howpublished = {Online research report, Dimensional Research},
  year         = {2026},
  url          = {https://research.dimensional.org/system-one-navigation/}
}

References.

  1. Khanna, M., Mao, Y., Jiang, H., Haresh, S., Schacklett, B., Batra, D., Clegg, A., Undersander, E., Chang, A. X., Savva, M. "Habitat Synthetic Scenes Dataset (HSSD-200): An Analysis of 3D Scene Scale and Realism Tradeoffs for ObjectGoal Navigation." CVPR 2024. https://arxiv.org/abs/2306.11290
  2. Savva, M. et al. "Habitat: A Platform for Embodied AI Research." ICCV 2019. https://arxiv.org/abs/1904.01201
  3. Szot, A. et al. "Habitat 2.0: Training Home Assistants to Rearrange their Habitat." NeurIPS 2021. https://arxiv.org/abs/2106.14405
  4. Anderson, P. et al. "On Evaluation of Embodied Navigation Agents." 2018. https://arxiv.org/abs/1807.06757
  5. Datta, S. et al. "Integrating Egocentric Localization for More Realistic Point-Goal Navigation Agents." CoRL 2020 (SoftSPL, Habitat Challenge). https://arxiv.org/abs/2009.03231
Open the full task explorer ↗