Index / Perplexity

Member of Technical Staff (Software Engineer, Multimodal)

Perplexity

Pay
$220k–$405k
Workplace
On-site
Location
San Francisco · California · United States
First seen
10 days ago
Last seen
27 minutes ago
Board
Ashby

Summary

We are hiring builders to define how people talk to, show things to, and hear from AI In 2026, we launched Computer, the defining product for the new era of agentic AI.

  • rust

Posting

We are hiring builders to define how people talk to, show things to, and hear from AI In 2026, we launched Computer, the defining product for the new era of agentic AI. We've scaled beyond the millions of people using Perplexity every day for research, shopping, investing and curiosity into a new paradigm of using AI to transform knowledge into action. The Multimodal team builds the experiences and infrastructure that move AI interaction beyond touch and text — realtime voice, vision, and the platform systems behind them. We own the full path from a user speaking into a device to an answer coming back: the realtime session infrastructure that connects clients to frontier audio models, the backend orchestration that routes, records, and supervises live sessions, and the SDK that powers voice and multimodal experiences across Perplexity's apps. As a backend engineer on Multimodal, you will design and scale the distributed systems that carry live voice sessions in production — and drive entirely new products at the intersection of voice, vision, and agents. WHY PERPLEXITY IS DIFFERENT - Craftsmanship. We build high quality, tasteful products targeting both the AI native and AI curious. - Ownership. You identify the problem, design the solution and ship it. - Entrepreneurship. We think like founders, act with urgency, and hustle to deliver for each other and our users. - Scholarship. Work among highly talented peers, pursuing knowledge and truth, upleveling ourselves, our teams, and our products. - Partnership. We amplify each others' strengths, break down silos, and give selflessly to help our colleagues deliver excellence. WHAT YOU'LL DO - Design, build, and scale the backend session-worker architecture that powers realtime voice: durable per-session workers, provider routing, and stateful streaming over gRPC. - Own distributed-systems problems end-to-end — session lifecycle, crash recovery, reconnection and replay, multi-region deployment, and graceful degradation under real production load. - Build provider-agnostic streaming protocols from our backend to the Rust SDK that powers voice across every client stack. - Drive new products and initiatives in voice and multimodal AI from problem definition through technical design, implementation, and launch. - Build the orchestration layer that lets live voice models delegate work to tools, agents, and long-running tasks — safely, asynchronously, and at scale. - Partner closely wi