Security Software Engineer - Python & AI Evaluation
Key Skills
Job Description
Help a top AI lab evaluate and improve large language models through security-focused coding tasks. Bring your software-engineering judgment and hands-on security experience to work involving vulnerabilities, exploit verification and security patches. This is a contracting engagement, with potential for a longer-term engagement. Remote candidates in the selected countries are elegible. What you will do Evaluate coding tasks involving software vulnerabilities, exploit verification and security patches. Create high-quality coding prompts and reference answers for benchmark-style problems. Evaluate model outputs for code generation, refactoring, debugging and implementation. Identify and document model failures, edge cases and reasoning gaps. Compare private language models with leading external models. Build or configure coding environments for evaluation and reinforcement learning. Follow detailed annotation and evaluation guidelines consistently. What you bring At least five years of professional software-development experience and strong Python skills. Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting or validating security patches. The ability to apply structured evaluation criteria and write clear technical feedback. Fluency in written and spoken English. Helpful, not required Professional code review, coding annotation, LLM/code evaluation or benchmark design. Knowledge of another programming language. Team leadership or mentoring experience.
Core Responsibilities
Evaluate AI-generated code on security tasks, including vulnerability analysis, exploit verification, and security patches, and create benchmark prompts and reference answers. Document model failures, compare models, and configure coding environments for evaluation and reinforcement learning.
Requirements
Requires at least five years of professional software-development experience, strong Python skills, and hands-on experience with vulnerability research, exploit verification, or security patches. Candidates must be fluent in written and spoken English and able to apply structured evaluation criteria and provide clear technical feedback.
About Braintrust
Industry: Technology, Information and Internet
Company size: 11-50 employees
Braintrust is revolutionizing hiring with Braintrust AIR, the world's first and only end-to-end AI recruiting platform. Trained with human insights and proprietary data, Braintrust AIR reduces time to hire from months to days, instantly matching you with pre-vetted qualified candidates, and conducting the first round phone screen for you. Trusted by hundreds of Fortune 1000 enterprises including Nestlé, Porsche, Atlassian, Goldman Sachs, and Nike, Braintrust AIR is making talent acquisition professionals 100x more effective and saving companies hundreds of thousands of dollars in recruiting costs.