What are the responsibilities and job description for the Researcher - Bayesian Deep Learning & Scalable Sequence Modeling position at NJF Global Holdings Ltd?
Quantitative Research & ML Engineering Organization
Hiring across four specialized teams for a client running fast, empirical experiments on around half a billion time series observations, figuring out what actually generates signal, and scaling the systems that support that research.
Culture and environment:
- Small, highly collaborative teams with zero bureaucratic overhead. You sit alongside peers who share context directly, iterate together daily, and debug side by side.
- The talent density is exceptionally high, with a strong contingent of former International Math Olympiad medalists, competitive programming finalists, and scientific researchers across the group.
- For research roles, the bar is first-principles mathematics first, machine learning second. You should be able to derive models, reason about measure-theoretic probability, and understand foundational statistical physics or dynamical systems before reaching for a standard neural framework.
- For engineering roles, the bar is first-principles computer science first, machine learning second. You need a deep understanding of memory hierarchies, cache coherence, compiler internals, and operating systems before worrying about high-level ML abstractions.
The firm owns and operates massive dedicated data center infrastructure. Researchers and engineers have direct access to abundant, modern GPU compute without queuing bottlenecks or resource rationing slowing down iteration speed.
The operating horizon is mid-frequency. The focus is on intraday and multi-day temporal structure rather than microsecond hardware races, which means research effort goes straight into predictive modeling, representation learning, and probabilistic guarantees rather than colocated FPGA plumbing.
This is an applied, proprietary research environment. If your primary career goal is writing academic papers and presenting at NeurIPS or ICML, this is not the place for you. The objective function is singular: generating the highest possible Sharpe ratio while deploying architectures that scale across substantial balance sheet capacity. If you want to push the absolute frontier of modern temporal modeling, Bayesian inference, and high-performance ML systems on real, noisy data with immediate feedback loops, you will thrive here.
Team 1: Sparse Architectures & Dynamic Routing
This team builds small, efficient Mixture of Experts (MoE) models that adapt to shifting regimes across fragmented markets.
Kinds of things you will work on:
- Running daily training experiments on novel routing algorithms and gating networks to see what handles noisy mid-frequency market regimes without collapsing.
- Testing sparse matrix kernels and memory layouts in CUDA/Triton to drop forward-pass latency down to the bare minimum.
- Figuring out why an expert allocation strategy fails on dynamic multi-hour order flow patterns and rapidly trying alternative balancing losses.
What you bring:
- Core foundation: First-principles math (optimization theory, linear algebra) and systems CS (hardware memory models, sparse data representations).
- Practical experience experimenting with conditional computation, sparse models, or deep learning systems.
- Comfort writing and optimizing GPU kernels on NVIDIA hardware.
- Intuition for mid-frequency market dynamics and non-stationary time series.
Team 2: Foundational Probabilistic Sequence Modeling
This group explores large-scale Bayesian sequence-to-sequence architectures and hybrid systems over historical datasets reaching hundreds of millions of events.
Kinds of things you will work on:
- Experimenting with hybrid formulations that fuse traditional structural priors and state space models with modern deep sequence backbones like SSMs, linear transformers, and continuous-time jump processes.
- Running empirical sweeps on variational inference, Laplace approximations, and stochastic gradient MCMC to evaluate calibration across multi-horizon, mid-frequency forecasts.
- Testing combinations of parametric market dynamics with non-parametric neural components to see where inductive biases outperform purely data-driven sequence learners.
- Measuring how uncertainty bounds hold up during extreme liquidity shocks and tuning loss functions to penalize uncalibrated confidence.
What you bring:
- Core foundation: Rigorous first-principles mathematics (measure-theoretic probability, Bayesian inference, functional analysis, stochastic calculus). Olympiad-level mathematical maturity is very welcome.
- Applied experience designing hybrid probabilistic architectures that blend classical time series methods or state space filters with deep sequence networks.
- Hands-on track record running multi-GPU distributed training runs in JAX or PyTorch across large internal cluster footprints.
- Clear understanding of modern sequence modeling literature and generative time series approaches.
Team 3: ML Performance, Kernel Engineering & Model Debugging
This group acts as the SWAT team for the research lab. When an experiment is inexplicably slow or silently blowing up, you find the root cause.
Kinds of things you will work on:
- Investigating why a researcher's simulation is running at 10% GPU utilization and rewriting the pipeline to make it fast across dedicated nodes.
- Digging through numerical errors, sudden NaN losses, broken JAX XLA compilation graphs, and silent PyTorch memory leaks.
- Profiling custom operators with Nsight Systems, writing custom Triton or CUDA kernels, and eliminating host-to-device synchronization bottlenecks.
- Building profiling harnesses so researchers can see instantly where their training loops are stalling.
What you bring:
- Core foundation: Hardcore first-principles computer science (operating systems, computer architecture, cache hierarchies, compilers). Competitive programming experience (ICPC/IOI) is a significant plus.
- Deep expertise in JAX internals (XLA tracing, compilation gotchas, sharding) and PyTorch runtimes.
- Fluency with hardware profiling tools like NVIDIA Nsight Compute and Nsight Systems.
- Sharp eye for numerical stability, floating-point arithmetic edge cases, and low-level systems execution in C or CUDA.
Team 4: Applied Production & Market Execution
This team takes the models that survive research validation and wires them directly into live execution loops across specific financial instruments.
Kinds of things you will work on:
- Experimenting with model variants across mid-frequency horizons, testing what works best for commodity calendar spreads versus equity dark pools, options surfaces, or CDS curves.
- Building deterministic C or Rust inference pathways that ingest live market feeds and output portfolio state transitions.
- Designing online monitoring systems to track live feature pipelines, asynchronous temporal alignment, and production model drift.
What you bring:
- Core foundation: Strong systems CS foundation (low-level memory management, deterministic concurrency, network socket programming).
- Systems programming experience in high-throughput production environments using C or Rust.
- Background in electronic trading, systematic execution, or financial engineering.
- Working familiarity with the mechanics of at least one major asset class (such as options pricing and Greeks, CDS roll basis).