Braintrust logo

Senior Application Security Engineer - AI Code Evaluation

Braintrust

Contractor, Full-time
5-10 years experience
Remote Only

Key Skills

Python
Application Security
Vulnerability Research
Exploit Verification
Security Patching
Software Development
Code Review
LLM Evaluation
Benchmark Design
Code Generation
Refactoring
Debugging
Technical Writing
Evaluation Criteria
Reinforcement Learning

Job Description

Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches. This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible. What you will do Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches. Create high-quality coding prompts and reference answers for benchmark-style problems. Evaluate model outputs for code generation, refactoring, debugging and implementation. Identify and document model failures, edge cases and reasoning gaps. Compare private language models with leading external models. Build or configure coding environments for evaluation and reinforcement learning. Follow detailed annotation and evaluation guidelines consistently. What you bring At least five years of professional software-development experience and strong Python skills. Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches. The ability to apply structured evaluation criteria and write clear technical feedback. Fluency in written and spoken English. Helpful, not required Professional code review, coding annotation, LLM/code evaluation or benchmark design. Knowledge of another programming language. Team leadership or mentoring experience.

Core Responsibilities

Evaluate AI model coding outputs on software vulnerabilities, exploit verification, and security patches, and create benchmark prompts with reference answers. Document model failures and compare private and external models while configuring coding environments and following evaluation guidelines.

Requirements

Requires at least five years of professional software-development experience, strong Python skills, and hands-on experience with vulnerability research, exploit verification, or security patches. Candidates must apply structured evaluation criteria, provide clear technical feedback, and be fluent in written and spoken English.

About Braintrust

Industry: Technology, Information and Internet

Company size: 11-50 employees

Braintrust is revolutionizing hiring with Braintrust AIR, the world's first and only end-to-end AI recruiting platform. Trained with human insights and proprietary data, Braintrust AIR reduces time to hire from months to days, instantly matching you with pre-vetted qualified candidates, and conducting the first round phone screen for you. Trusted by hundreds of Fortune 1000 enterprises including Nestlé, Porsche, Atlassian, Goldman Sachs, and Nike, Braintrust AIR is making talent acquisition professionals 100x more effective and saving companies hundreds of thousands of dollars in recruiting costs.

Added Today