Reinforcement Learning
Hierarchical & meta multi‑agent RL, PPO / SAC / A2C baselines, and AlphaZero‑style MCTS agents applied to power‑grid topology control (L2RPN) and physics‑driven combat environments.
Hi, I'm
I work on
Machine Learning Engineer & independent researcher building reinforcement‑learning agents, motion & pose foundation models, and vision systems — currently exploring language‑guided RL and large‑scale mocap pretraining.
Get to know
I'm an ML Engineer focused on building production‑grade ML and AI systems that solve real‑world problems.
Currently at Home of Performance (HOP), I design and build LLM‑based systems and agentic workflows for clients, from architecture through deployment.
My research background spans reinforcement learning, power grid control, LLMs & SLMs, motion capture, and game AI — I enjoy working across both ends, from research to production.
What I work on
Hierarchical & meta multi‑agent RL, PPO / SAC / A2C baselines, and AlphaZero‑style MCTS agents applied to power‑grid topology control (L2RPN) and physics‑driven combat environments.
GPT‑ and Mamba‑style sequence models pretrained on large‑scale mocap data for human motion generation and prediction.
Vision Transformers for multi‑cancer & lung‑cancer classification, segmentation for autonomous driving, and skull‑stripping pipelines.
Small‑LM coding agents, DSPy‑powered synthetic data generation, and low‑resource language modelling for Tamil.
Graph neural networks, neural‑guided A* search and hierarchical/meta RL for grid topology optimization and congestion management on the L2RPN benchmark.
Selected work
Foundation model for human motion, pretrained on large‑scale mocap data spanning sports, combat, and tactical movement.
Generative Pre‑trained State Machine trained on large‑scale mocap data to predict and generate human motion.
Hierarchical multi‑agent reinforcement learning for power grid topology control on the L2RPN benchmark.
AlphaZero‑based topology optimization agent for power grid congestion management using Monte Carlo Tree Search.
Neural‑guided A* search for power grid topology optimization.
Claude Code‑inspired coding agent powered by small language models instead of large LLMs.
A study of Large Language Models for the Tamil language — exploring low‑resource language modelling.
A unified tool to generate synthetic text and sensor data using LLMs with the help of DSPy.
Vision Transformer architecture for multi‑cancer image classification.
Deep reinforcement learning agents built with Kolmogorov–Arnold Networks (KAN).
Toolbox
Get in touch
Open to research collaborations, ML engineering roles, and interesting problems in reinforcement learning or computer vision. Reach out.