FOUNDING ROLE BRIEF / SIMLABS RESEARCH / 01 OF 05

Lead AI Researcher

Own two connected tracks: the applied intelligence that makes the Terminal dependable now, and the spatial research that can make venue-state learning, simulation and physical evaluation useful later.

ROLE BRIEFLOCATION + TERMS TO BE PUBLISHEDEXPRESS INTEREST

THE ROLE / WHAT YOU OWN

About this role.

The immediate problem is applied: make the current Terminal software choose approved tools, ground every answer in authoritative venue state, recover safely in noise and run within a deployment profile we can defend. You will own the evaluation harness, model adaptation and local-versus-cloud decisions that turn a demo into a pilot-ready system.

The research agenda is separate and real: representations of changing venue state, correspondence between simulated and physical episodes, and action-and-outcome records for spatial-model evaluation. Useful work may become shipped code, a benchmark, a dataset schema or a paper. We want results other researchers can inspect, reproduce and challenge.

THE RAMP / START WITH WHAT EXISTS

Your first 90 days.

Every role starts with real work on the current software prototype and the first pilot plan. We will agree the milestones together and measure the result.

FIRST 30 DAYS

Publish the first reproducible baseline for the current software demo: grounding, refusal, noisy-audio robustness, latency and multilingual coverage.

BY DAY 60

Choose and validate the first pilot model profile, including what runs locally, what may use a provider and what deterministic services must own.

BY DAY 90

Write the simLabs research agenda and release one inspectable artifact: a venue-state benchmark, experience-record schema or simulation-evaluation harness.

RESPONSIBILITIES / THE WORK

What you will do.

  1. 01

    Own the model roadmap: what we use, adapt, train or keep deterministic, and what runs locally or through a provider

  2. 02

    Build the evaluation harness for grounded tool use, refusal behavior, noisy-audio robustness, latency and multilingual coverage

  3. 03

    Adapt or fine-tune models only when a measured baseline shows that it improves the service

  4. 04

    Define the spatial research programme around venue-state representation, simulation-to-reality evaluation and action-and-outcome episode schemas

  5. 05

    Turn demo and future pilot failures into versioned evaluation scenarios with provenance

  6. 06

    Define training and evaluation dataset standards with the data team under a documented lawful basis and separate, purpose-specific rights

  7. 07

    Publish protocols, benchmarks and results when review, evidence and data rights permit

TECH / THE STACK

What you will work with.

The current Terminal is a local software demo. Deployment profiles may combine local, edge and cloud components; every provider must earn its place on latency, privacy, cost and reliability.

Modeling

  • Python and PyTorch
  • Open-weight model families
  • Fine-tuning, LoRA and adapters
  • Distillation and quantization

Serving and edge

  • Local inference on development and target Terminal hardware
  • Selected cloud inference where justified
  • CUDA, TensorRT and ONNX Runtime
  • Speech recognition and speech synthesis pipelines

Evaluation

  • Reproducible evaluation harnesses
  • Human-in-the-loop review flows
  • Regression suites built from demo scenarios and future pilot failures

Research artifacts

  • Venue-state benchmarks
  • Simulation and physical episode comparison
  • Dataset and model cards
  • Reproducible releases

QUALIFICATIONS / THE BAR

What you bring.

Must have

  • 6+ years in applied machine learning with deep, hands-on transformer experience
  • You have taken fine-tuned or distilled models into production and lived with the consequences
  • Evaluation-first instincts: you distrust a demo until the harness agrees
  • Strong Python and PyTorch engineering; you turn exploration into reproducible work that survives review and release
  • Experience with voice or conversational systems in the real world
  • Comfortable owning a research agenda inside a small, senior team

Bonus points

  • Spatial perception, embodied AI or world-model research
  • Simulation, robotics or sim-to-real evaluation
  • Edge and on-device inference experience
  • Retrieval and grounding architectures for factual answers
  • Multilingual model work
  • Published research or open-source benchmark contributions

THE CULTURE / DAY TO DAY

How we work.

EVIDENCE OVER OPINION

Prototypes, evaluation harnesses, non-team tests and pilot measurements settle debates. The best argument is inspectable evidence.

PROTOTYPE TO PILOT

Work moves from the current software demo to hardware tests and then to a bounded venue pilot. Each stage is labeled, measured and allowed to fail honestly.

SMALL AND SENIOR

No layers, no committees. Five hires, two founders, one room when it matters.

DESIGNED FOR REAL FLOORS

We test for noise, glare, accessibility, latency and failure before a venue depends on the system.

WHAT WE OFFER / TERMS IN WRITING

Know the terms before you decide.

Cash and equity bands, operating location, employment jurisdiction and work-authorization policy are not yet published. Before an interview process advances, we will supply them and separate cash, equity, vesting and any profit-sharing terms in writing.

  • A written offer that separates cash, equity, vesting and any profit-sharing terms
  • Role scope and decision rights agreed before the start date
  • Operating location, employment route, work authorization and travel expectations confirmed before interviews advance
  • Direct work with both founders and senior ownership of a defined product layer
  • A working pattern agreed for the role, including field research and pilot travel where needed
  • A hardware and tools budget agreed before the start date

APPLY / SIMLABS RESEARCH

Tell us what you would build first.

Twelve months from now, the first Terminal pilot should run on an intelligence profile you measured, and simLabs should have released at least one spatial research artifact you are proud to put your name on. Send a short note with a repo, model card, benchmark, dataset paper, evaluation harness or a war story about a model that failed interestingly. A conversation with a founder follows within days.

Express interest in Lead AI Researcher