Post Job Free
Sign in

Artificial Intelligence Engineer / Architect

Company:
Predii
Location:
San Mateo, CA
Posted:
September 01, 2026
Apply

Description:

AI Engineer / Architect\n Senior Agentic AI & Applied ML\n US-based — CA preferred, open to West Coast + remote.

Hybrid-friendly.\n ABOUT PREDII\n n Predii builds the intelligence layer that runs the automotive service and parts industry.

Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence — powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.\n We're small, fast, and allergic to red tape.

No 12-layer approval chains, no work that disappears into a backlog forever.

If you build something here, it ships — and real customers use it.

Learn more at RESEARCH\n n We do real research, not just integration.

We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content.

We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead.

And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money.

All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.THE VIBE\n n We need an AI Engineer/Architect who wants more than a model to fine-tune — someone ready to own agentic AI products end-to-end and shape the platform they run on.

This is real ownership, not busywork.

You'll touch:\n\n Agentic products across dealership operations — technician diagnostics, service advisor recommendations, pricing/inventory intelligence\n A shared architecture that supports multiple agentic products, not one-off builds per use case\n Inference performance and cost — latency and cost-per-token as design inputs, not afterthought metrics\n Evaluation frameworks tuned to what "correct" means for each application\n Safety guardrails and escalation paths for decisions with real diagnostic and financial impact\n Integration reality across varying DMS/shop management systems\n\n Expect to shape architecture and set technical direction, not just execute someone else's roadmap.WHAT YOU'LL ACTUALLY DO\n n\n Own Agentic Products End-to-End — from concept to production, across technician diagnostics, service advisor recommendations, and pricing/inventory intelligence.\n Build the AI Layer on Shop/Dealer Systems — agentic applications that reason over past repair orders, monitor DMS activity in real time, and act on unstructured pricing/sourcing data.\n Design a Shared Platform — one architecture that supports multiple agentic products, not a pile of bespoke builds per use case.\n Build Evaluation Frameworks — tailored to each application (diagnostic accuracy, recommendation relevance, pricing correctness), both offline and in production.\n Architect Inference for Scale — high-throughput, low-latency serving on open-weight models; own cost-per-token and latency-per-request as design inputs, not just dashboards to watch.\n Define Safety & Escalation — guardrails and human-in-the-loop paths for high-stakes recommendations with diagnostic or financial impact.\n Work the Real Integration Surface — varying DMS/shop management system integrations as a core design constraint, not an afterthought.\nYOUR TOOLKIT\n n\n Models & Serving: Open-weight LLMs (Llama, Mistral, Qwen, or similar), vLLM/TGI/TensorRT-LLM, GPU inference optimization\n Agentic Frameworks: LangGraph, LangChain, custom agent/orchestration stacks, tool-use & function calling\n Evaluation & Observability: Offline eval harnesses, production monitoring (accuracy, relevance, drift), LLM-as-judge, tracing (LangSmith or similar)\n Data & Retrieval: RAG pipelines, vector databases, structured + unstructured data reasoning over repair orders and DMS records\n Languages & Systems: Python, distributed systems fundamentals, API design for real-time DMS integrations\n Cloud & Infra: Azure/GCP/AWS, containers/Kubernetes, cost and latency instrumentation\n Safety: Guardrails, escalation/human-in-the-loop design, risk classification for high-stakes outputs\nWHAT YOU BRING\n n\n 5+ years building and shipping production ML/AI systems, with real users depending on the output — not just research or labs.\n Hands-on experience building agentic or LLM-powered applications end-to-end, from prototype to production.\n Track record architecting shared platforms/services that support multiple product use cases, not single-purpose builds.\n Practical experience with evaluation frameworks for AI systems — designing metrics that map to business correctness, not just model benchmarks.\n Experience optimizing inference for throughput/latency/cost on open-weight models in production.\n Comfort designing safety guardrails and escalation logic for high-stakes, high-consequence recommendations.\n Experience integrating with messy, heterogeneous third-party systems (APIs, DMS, ERPs, or similar) as a given constraint.\n A self-starter mindset — comfortable owning ambiguity and working independently across a distributed US–India team.\nBONUS POINTS\n n\n Background in automotive, dealership, or other data-heavy vertical platforms.\n Experience with diagnostic, recommendation, or pricing systems where wrong answers have real financial/safety stakes.\n Prior work fine-tuning or serving open-weight models at scale (not just calling a hosted API).\n Startup or small-team energy — you've owned a product, not just a component.\n Technical leadership or mentoring experience.\nHOW WE ROLL\n n\n Own it — flag issues early, make the call, follow through.\n Be proactive — don't wait to be told; spot the risk, bring the fix.\n Share the load — quality and reliability are everyone's job, not just yours.\n Make it count — your work ships to production and touches real customers, real fast.\n\n How To Apply\n\n Email: \n\n

Apply