We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
Remote

Reinforcement Learning Engineer (Cybersecurity)

Bugcrowd
$176,400 - $242,550
United States
Sep 11, 2026

Founded in 2012, Bugcrowd is the preemptive security platform that unifies exposure discovery and assessment, offensive testing, and intelligence shaped by AI and human insight to help organizations avoid, discover, and validate real-world risk. Bugcrowd helps security teams move faster by identifying the exposures that matter most so they can act first and stay ahead of attackers. By combining the power of humans and AI, teams can preempt attack paths and prevent breaches. Based in San Francisco and New Hampshire, Bugcrowd is supported by General Catalyst, Rally Ventures, Costanoa Ventures, and others. Visit www.bugcrowd.com.

Job Summary

The Bugcrowd RL and Reasoning Team focuses on pushing the boundaries of autonomous cybersecurity by building authentic, verifiable reinforcement learning environments for world-leading foundational AI companies. As a Reinforcement Learning Engineer specializing in Reinforcement Learning from Verifiable Rewards (RLVR), you will design and scale automated verification pipelines that transform real-world software vulnerabilities into deterministic reward functions. In this role, you will bridge the gap between low-level security analysis and modern LLM reasoning models, engineering environments where AI agents learn to discover, exploit, and remediate software vulnerabilities with mathematical certainty. Instead of relying on subjective human feedback, your work directly powers the rigorous, verifiable reward signals that teach next-generation frontier AI models how to master complex cybersecurity domain logic. You will work at the intersection of fuzzing, dynamic program analysis, system exploitation, and scalable ML infrastructure to shape the safety and offensive/defensive capabilities of future artificial intelligence.

Essential Duties and Responsibilities



  • Design, build, and deploy high-throughput RLVR (Reinforcement Learning from Verifiable Rewards) environments that evaluate LLM action sequences against deterministic execution outcomes.
  • Develop automated test harnesses, sandboxes, and verification engines that convert complex vulnerability research (e.g., memory corruption, web security, logic bugs) into binary pass/fail reward signals.
  • Integrate Bugcrowd's Mayhem automated analysis platform and real-world vulnerability feeds into continuous, scalable RL environment generation pipelines.
  • Architect safe, isolated, and highly reproducible execution environments (using Docker, BuildKit, or Nix) capable of running thousands of simultaneous agent-driven exploitation and patching trajectories.
  • Collaborate directly with researchers at frontier AI labs including Anthropic, OpenAI, and Cohere to define standard benchmark formats, observation spaces, and verifiable evaluation metrics for cybersecurity tasks.
  • Implement precise telemetry, ground-truth verification algorithms, and trajectory logging to analyze agent reasoning paths and prevent reward hacking or false positives.
  • Build low-level instrumentation and debugging tools to monitor memory states, process executions, and network behaviors during agent interaction cycles.
  • Optimize infrastructure performance and environment reset latency to support massive-scale parallel sampling and distributed RL training workflows.
  • Benchmark and evaluate frontier AI model performance across diverse offensive and defensive security challenges, such as automated fuzzing, exploit payload generation, and patch validation.


Education, Experience, Knowledge, Skills, and Abilities



  • Understanding of RL training workflows used by modern LLM systems, specifically execution-based feedback or Reinforcement Learning from Verifiable Rewards (RLVR).
  • Proficiency developing applications in Python and low-level systems programming in C, with Rust experience being a strong plus.
  • Solid understanding of software vulnerabilities, binary exploitation, fuzzing methodologies, or program analysis.
  • Experience with DevOps pipelines (e.g., GitHub Actions), reproducible builds (Docker, BuildKit, Nix), and comfort working with Linux systems and low-level debugging.
  • Experience working with or building benchmark environments (e.g., CTFs, SWE-bench, security challenges, or execution sandboxes).


Preferred Experience

  • Experience designing custom reward functions, ground-truth verifiers, or automated grading engines for AI safety and reasoning models.
  • Background in low-level program analysis tools, sanitizers (e.g., ASan/MSan), compiler instrumentation, or automated exploit generation tools.
  • Proven track record of participating in or developing competitive cybersecurity benchmarks, CTFs, or open-source AI evaluation frameworks.


Working Conditions and Physical Requirements

The ideal candidate must be able to complete all physical requirements of the job with or without reasonable accommodation.

Sitting and / or standing - Must be able to remain in a stationary position 50% of the time

Carrying and / or lifting - Must be able to carry / move laptop as needed throughout the work day.

Environment - remote, work-from-home 100% of the time.

Pay Range Disclosure

At Bugcrowd, we strive for fairness, equality and to create an environment that allows our people to perform at their very best. Our compensation philosophy is to foster a collaborative community that rewards, attracts and retains the best possible talent. The provided salary details are based on US national averages and we retain the flexibility to tailor to the needs of the business.

The national estimate for the current base range for the position of $176,400 - $242,550.

This position may also be eligible to participate in a discretionary bonus program or commission plan, subject to the rules governing the program, whereby an award, if any, depends on various factors, including, without limitation, individual and organizational performance.

Culture



  • At Bugcrowd, we understand that diversity in the workplace is vital to a company's success and growth. We strive to make sure that people are included and have a sense of being part of making Bugcrowd not only a great product but a great place to work.
  • We regularly hear from both customers and researchers that Bugcrowd feels like a family, and we strive to maintain that internally as well.
  • Our team consists of a broad range of people: musicians, adventure sports junkies, nature lovers, parents, cereal enthusiasts, night owls, cyclists, artists-you get the point.


At Bugcrowd, we are solving security threats and vulnerabilities that are relevant to everyone, therefore we believe solving these problems takes all kinds of backgrounds. We value the perspectives and experiences people from underrepresented backgrounds bring.

Disclaimer

This position has access to highly confidential, sensitive information relating to the technologies of Bugcrowd. It is essential that the applicant possess the requisite integrity to maintain the information in the strictest confidence.

The company is authorized to obtain background checks for employment purposes under state and federal law. Background checks will be conducted for positions that involve access to confidential or proprietary information (including trade secrets).

Background checks may include Social Security verification, prior employment verification, personal and professional references, educational verification, and criminal history. Applicants with conviction histories will not be excluded from consideration to the extent required by law.

Any personal data you submit in connection with your application will be processed in compliance with Bugcrowd's Privacy Policy, which you may review here:https://www.bugcrowd.com/privacy.

Equal Employment Opportunity:

Bugcrowd is EOE, Disability/Age Employer.

Individuals seeking employment at Bugcrowd are considered without regards to race, color, religion, national origin, age, sex, marital status, ancestry, physical or mental disability, veteran status, gender identity, or sexual orientation.

Bugcrowd is committed to the full inclusion of all qualified individuals. In keeping with our commitment, Bugcrowd will take the steps to assure that people with disabilities are provided reasonable accommodations. Accordingly, if reasonable accommodation is required to fully participate in the job application or interview process, to perform the essential functions of the position, and/or to receive all other benefits and privileges of employment, please contact HR at ADA at bugcrowd.com.

Apply at:https://www.bugcrowd.com/about/careers/

Applied = 0

(web-9db6c7984-zzklj)