Staff Software Engineer, Frontier Security Team
Snowflake
- Pay
- $236k–$295k
- Workplace
- On-site
- Location
- US-CA-Menlo Park · Menlo Park · California · United States
- First seen
- 2 hours ago
- Last seen
- 2 hours ago
- Board
- Ashby
Summary
At Snowflake, we are powering the era of the agentic enterprise.
Posting
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are hiring a Staff Software Engineer for our Frontier Security AI team. Snowflake's Frontier Security AI teams develop production-grade LLM applications, intelligent agents, AI infrastructure, and evaluation systems for enterprise customers — products that must meet a high bar for quality, security, reliability, and efficiency while operating over sensitive data at large scale. In this role, you will lead the design and development of our Agentic Harness and agent evaluation platform, working across product, infrastructure, applied AI, security, and modeling teams to take new capabilities from prototype to dependable customer value. AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL: - Architect and build the Agentic Harness that executes complex, multi-step AI workflows across models, tools, data, and services. - Design stable interfaces for tool execution, context construction, state management, memory, permissions, retries, fallbacks, and human review. - Own agent quality end to end by building evaluation harnesses, representative datasets, automated graders, experiment pipelines, and release gates. - Convert ambiguous reports such as "the agent feels worse" into measurable failure modes, reproducible tests, and durable fixes. - Analyze production agent trajectories to identify failures in reasoning, retrieval, tool use, context, orchestration, and application code. - Close the loop between production incidents, root-cause analysis, evaluation coverage, and regression prevention. - Develop offline and online measurements for task completion, correctness, groundedness, safety, latency, reliability, and cost. - Build simulation and replay infrastructure for golden-set tests, adversarial scenarios, model comparisons, and large-scale exper