Jobgether logo

SWE-Bench AI Task Auditor - Freelance AI Trainer Project

Jobgether

Contractor
5-10 years experience
Remote Only

$60 per hour

Key Skills

Software engineering
Codebase navigation
Troubleshooting
Test failure analysis
Logic error identification
Technical validation
SWE-Bench
Problem-solving
Analytical skills
Technical documentation
AI model evaluation
Code review

Job Description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SWE-Bench AI Task Auditor - Freelance AI Trainer Project based in Netherlands.

This freelance opportunity is designed for experienced software engineers who want to contribute their technical expertise to the development and evaluation of AI systems.
You will review software engineering tasks used to train and assess advanced AI models.
Your work will focus on ensuring tasks are technically accurate, realistic, reproducible, and appropriately tested.
You will investigate codebase integrations, test failures, logic issues, and other technical challenges.
The role offers flexibility through a fully remote, project-based working model.
You will apply practical software engineering judgment rather than simply following predefined checks.
Your feedback will directly help improve the quality and reliability of AI training workflows.

\n


Accountabilities
  • Evaluate software engineering tasks for technical accuracy, realism, solvability, reproducibility, and alignment with SWE-Bench standards.

  • Review task codebases, integrations, tests, and evaluation criteria to identify potential technical weaknesses.

  • Rigorously test and troubleshoot complex technical scenarios to determine whether tasks function as intended.

  • Investigate codebase integration problems, test failures, logic errors, and other implementation issues.

  • Provide clear, precise, and actionable feedback that enables task creators to correct identified problems.

  • Apply professional software engineering judgment to assess whether tasks reflect realistic development scenarios.

  • Help maintain a high standard of technical rigor and accuracy across AI training and evaluation workflows.

Requirements

  • Demonstrable professional experience in software engineering, including experience navigating complex codebases and developing real-world applications.

  • Strong knowledge of software engineering principles and familiarity with SWE-Bench-style tasks and evaluation workflows.

  • Strong analytical and problem-solving abilities, with the capacity to investigate complex technical scenarios systematically.

  • Experience troubleshooting code, diagnosing test failures, and identifying underlying logic or integration issues.

  • Ability to evaluate technical work objectively and communicate findings clearly and constructively.

  • Strong attention to detail and a rigorous approach to technical validation.

  • Ability to work independently and manage project-based assignments effectively.

  • Deep expertise in one relevant software engineering specialty is sufficient; expertise across multiple domains is not required.

  • A secure computer and reliable, high-speed internet connection suitable for remote technical work.

Benefits

  • Fully remote freelance contract.

  • Flexible project-based working environment.

  • Opportunity to contribute directly to the training and evaluation of AI systems.

  • Ability to apply real-world software engineering expertise to challenging technical tasks.

  • Compensation of $60/hour, with the final rate determined based on experience, expertise, and geographic location.

  • No company-sponsored health insurance, PTO, or other employee benefits, as this is a freelance contractor position.


\n

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

 Why Apply Through Jobgether? 

 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

 

 

#LI-CL1

Core Responsibilities

You will evaluate and troubleshoot software engineering tasks to ensure they meet technical standards for AI model training and assessment. Your role involves providing actionable feedback to task creators and maintaining high standards of technical rigor across evaluation workflows.

Requirements

Candidates must have demonstrable professional experience in software engineering and a strong ability to navigate complex codebases. You should possess deep expertise in a software engineering specialty and the ability to work independently on project-based assignments.

Benefits

  • Fully remote freelance contract
  • Flexible project-based working environment

About Jobgether

Industry: Internet Marketplace Platforms

Company size: 11-50 employees

Jobgether is a career navigation platform built for senior professionals competing in the remote job market. Hiring systems were designed for volume and keyword matching, not for 15 or 20 years of nonlinear experience. That structural mismatch is why strong profiles get filtered out before a human ever sees them. Our platform diagnoses where a search is breaking, corrects how experience is positioned, and connects professionals to the companies where their background creates real value. The goal is not more applications. It is the right visibility in the right places. Our mission is to ensure no senior professional remains invisible in the global remote market, not because they lack the skills, but because the system failed to read them correctly.

Added Today