Founding Data Engineer
San Francisco
About the company
Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them.
Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve.
About the Role
Braintrust is growing quickly across enterprise and self-serve customers, and our internal systems need to scale with us. You will establish our data foundation, make architectural decisions, work directly with our executive team, and help shape the data function as it grows. This is a founding, hands on role. You will decide what good looks like, and be able to build it.
The core of the job is analytics engineering. That means defining the business entities and metrics the company runs on, modeling them so they can be trusted, and making them accessible without a data person in the loop. Around that core, you will own the ingestion pipelines that feed it, the architecture decisions that hold it together, and the reliability of the whole path.
You'll partner with stakeholders across the business, and directly with executives. You will work with Product and Engineering on usage data and event schemas, with Finance on revenue and consumption reporting, and with Sales, RevOps, and Marketing on the pipeline and customer data they operate from.
What You Will Do
Define the business entities, metrics, and semantics that Product, Finance, and GTM all build on
Design and build the core data models in SQL and dbt
Build and own ingestion from product events, CRM, and other business systems, including vendor APIs, webhooks, and sync gaps
Make architectural decisions by determining how storage, orchestration, and transformation evolve from here
Stand up self-serve analytics that includes building semantic models and trusted metrics that stakeholders can query without filing a ticket
Partner with Engineering and Product on usage data and event schemas as the product surface grows
Treat governance and reliability as part of the data model
Use AI to accelerate your own build, and keep the foundation clean and well-described enough that AI tooling on top of it has something reliable to stand on.
Where You’ll Work in the Stack:
Ingestion: Product events, CRM, and other business data sources
Modeling: SQL, dbt, business entities, and shared metric definitions
Storage and orchestration: The current environment includes Snowflake; you’ll help determine how the architecture evolves
Analytics: Semantic models and self-serve access to trusted metrics
Governance and reliability: Data quality, access controls, lineage, monitoring, and debugging
About You
You've defined business entities and metrics, and built the models behind them
Strong command of SQL and dbt, and experience with a cloud warehouse (Snowflake or equivalent) and an orchestrator (Dagster, Airflow, or similar)
Enough engineering depth to own pipelines end to end. You can write and maintain production code, not only transformation models.
Clear judgment about data modeling, schema evolution, contracts, and lineage
Strong debugging instincts across a multi-system path, and the patience to trace a wrong number to its source
Ability to work directly with executives and functional leaders: clarifying vague requests, explaining how the data works, and negotiating priorities and tradeoffs
Comfort operating as the only data hire, balancing foundational architecture against urgent business needs
Bonus Points
You've built a data foundation from scratch, rather than inheriting a mature one
Startup experience, especially as an early or first data hire
Experience across multiple business functions that includes Product, Finance, and GTM
Familiarity with usage-based pricing, consumption models, or PLG-plus-sales-assisted funnels
Experience with our stack, or close equivalents: Snowflake, dbt, Dagster, AWS Batch/Lambda, Fivetran, Terraform, GitHub Actions, Soda, or semantic-layer tools like Omni or Looker
You've built or used agents for internal workflows and analysis
Braintrust user :)
What Success Looks Like
Braintrust has a governed set of core entities and metrics that Product, Finance, and GTM all build on
Product usage data is reliable enough that people act on it without checking it against something else first
Stakeholders can answer their own questions through trusted, self-serve models
Every automated decision and every metric has a traceable path back to its source
The foundation is documented and maintainable. Simple enough for a startup team to run solo, structured enough to hand to a larger data team as it grows.
Benefits include
Medical, dental, and vision insurance
Daily lunch, snacks, and beverages
Flexible time off
Competitive salary and equity
Wifi & cellphone stipend
Equal opportunity
Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.