Portrait of Rohit AnanthanπŸ‘‹
Open to Data Science & AI roles

Hi, I'm Rohit

Data Scientist & AI Engineer

I spend my days turning messy, real-world data into AI systems people actually use β€” multi-agent pipelines for healthcare, fraud detection watching 5M+ transactions a day, and LLM apps that save teams real hours. Currently building with Python, GCP, Azure & GPT-4o.

4+ yrs
in production ML
10+
ML projects shipped
Springer
published research
2Γ—
AI certifications
01.

About Me

Data Scientist with 4+ years building production ML systems across healthcare, e-commerce, non-profit, and enterprise domains. Specializing in multi-agent AI systems, LLM-powered applications, product analytics, real-time anomaly detection, and scalable data engineering pipelines using Python, PySpark, Azure, GCP, and AWS.

With a background spanning Engineering (B.Tech from IIIT Chennai) to Information Systems (MS from University of Maryland β€” Smith School of Business), I bring a rare blend of engineering rigor and business context to every data problem.

PythonGCP Vertex AILLMs / RAGMLOpsSQLTableau
0+
Years Experience
0+
ML Projects Deployed
0%
Avg. Efficiency Gain
0+
LinkedIn Connections
πŸ“
United States
Open to Remote & Hybrid Roles
Available
rohit@portfolio: ~/profile.py
class DataScientist:
    def __init__(self):
        self.name      = "Rohit Ananthan"
        self.focus     = ["GenAI & RAG", "MLOps", "Graph ML", "Product Analytics"]
        self.stack     = {"lang": "Python", "cloud": ["GCP", "AWS"], "llm": "GPT-4o"}
        self.education = ["MS InfoSys @ UMD", "B.Tech @ IIIT Chennai"]

    def mission(self) -> Impact:
        # raw data in β†’ business impact out
        return bridge(technology, business_value)  # converged βœ“
02.

Experience

AI Engineer

Current

Vdart Inc.

Contract Β· Remote

Jun 2026 – Present
3 mos
  • β–ΈBuilding multi-agent AI systems that automate document-heavy enterprise workflows for a healthcare client (specifics under NDA)
  • β–ΈWorking across retrieval-grounded evaluation, agent orchestration, and full-stack delivery β€” with AI coding agents as part of the development workflow
Multi-Agent AIRAGAzureNode.jsReactAI-Assisted Delivery

Data Scientist Consultant

Invision Global Tech Inc

Full-time Β· United States

Feb 2026 – Jun 2026
5 mos
MLPythonData ScienceConsulting

Data Scientist

Community Dreams Foundation

Full-time Β· Remote

Feb 2025 – Feb 2026
1 yr
  • β–ΈBuilt an AI-powered Legal & Compliance Assistant using GPT-4o with a RAG pipeline on LangChain, Pinecone, and ANN search β€” cutting manual review time by 40%
  • β–ΈDeployed end-to-end ML pipelines on GCP Vertex AI and Dataproc with automated hyperparameter tuning via MLflow β€” improving demand forecasting accuracy by 15%
  • β–ΈImplemented MLOps workflows using Vertex AI, Cloud Build, and GitHub Actions for automated model versioning, drift detection, and CI/CD across environments
  • β–ΈBuilt a real-time fraud and anomaly detection system using Pub/Sub, Dataflow (Apache Beam), and XGBoost β€” reducing undetected fraud by 20%
  • β–ΈEngineered scalable data pipelines with Dataflow, Composer (Airflow), and BigQuery β€” cutting ETL latency by 60%
  • β–ΈDeveloped LLM-based apps for document summarization and Q&A using OpenAI APIs and Vertex Matching Engine for semantic search
  • β–ΈCreated predictive donor churn models with TensorFlow and Scikit-learn, deployed on Vertex AI with Looker Studio dashboards
GPT-4oLangChainGCPVertex AIXGBoostTensorFlowBigQueryAirflowMLflow

Financial Analyst

The Premiere Group

Full-time Β· Columbia, MO Β· On-site

Sep 2025 – Oct 2025
2 mos
Financial AnalysisData Analytics

Technical Consultant – Course Renewal Automation

University of Maryland – Extended Studies

Internship Β· College Park, MD Β· Remote

Jan 2024 – Dec 2024
1 yr
SalesforceAutomationMicrosoft Teams Planner

Graduate Assistant

University of Maryland

Part-time Β· College Park, MD Β· On-site

Jan 2024 – Dec 2024
1 yr
Data DigitizationCMSResearch

Data Scientist

Kameleon Technologies

Full-time Β· Chennai, India

May 2021 – Jul 2023
2 yrs 2 mos
  • β–ΈEngineered a real-time fraud detection pipeline using Neo4j graph database and XGBoost β€” processing 5M+ daily transactions with an 18% reduction in false positives
  • β–ΈBuilt scalable ETL/ELT pipelines on PySpark and AWS (S3, Glue, EMR) to process terabytes of financial data β€” reducing pipeline processing time by 35%
  • β–ΈDeveloped customer segmentation and churn prediction models using ensemble methods β€” driving a 12–15% improvement in customer retention across key segments
  • β–ΈEstablished MLOps practices with MLflow experiment tracking, SageMaker model registry, and automated retraining workflows for production model governance
  • β–ΈDelivered executive-facing Tableau dashboards for transaction monitoring, KPI tracking, and fraud trend analysis β€” adopted across operations and risk teams
Neo4jXGBoostPySparkAWSSageMakerMLflowTableauFraud Detection
03.

Skills

⚑

Languages & Libraries

PythonSQLRPySparkMATLABPandasNumPyPyTorchTensorFlowScikit-learn
🧠

ML & AI

Machine LearningDeep LearningMulti-Agent SystemsNLPLLMsGPT-4oRAGXGBoostGenAIA/B TestingCausal Inference
πŸ”§

Data Engineering

ETL / ELTApache AirflowApache BeamDataflowPub/SubDBTMLflowCI/CDGitHub Actions
☁️

Cloud & Infra

GCP Vertex AIAWS SageMakerAzureBigQuerySnowflakeDataprocCloud BuildDocker
πŸ“Š

BI & Visualization

TableauPower BILooker StudioData Storytelling
πŸ—„οΈ

Databases & Search

Neo4jAzure AI SearchMongoDBPineconeLangChainLlamaIndexVector DBs
Pythonβ—†
SQLβ—†
Rβ—†
PySparkβ—†
MATLABβ—†
Pandasβ—†
NumPyβ—†
PyTorchβ—†
TensorFlowβ—†
Scikit-learnβ—†
Machine Learningβ—†
Deep Learningβ—†
Multi-Agent Systemsβ—†
NLPβ—†
LLMsβ—†
GPT-4oβ—†
RAGβ—†
XGBoostβ—†
GenAIβ—†
A/B Testingβ—†
Causal Inferenceβ—†
ETL / ELTβ—†
Apache Airflowβ—†
Apache Beamβ—†
Dataflowβ—†
Pub/Subβ—†
DBTβ—†
MLflowβ—†
CI/CDβ—†
GitHub Actionsβ—†
GCP Vertex AIβ—†
AWS SageMakerβ—†
Azureβ—†
BigQueryβ—†
Snowflakeβ—†
Dataprocβ—†
Cloud Buildβ—†
Dockerβ—†
Tableauβ—†
Power BIβ—†
Looker Studioβ—†
Data Storytellingβ—†
Neo4jβ—†
Azure AI Searchβ—†
MongoDBβ—†
Pineconeβ—†
LangChainβ—†
LlamaIndexβ—†
Vector DBsβ—†
Pythonβ—†
SQLβ—†
Rβ—†
PySparkβ—†
MATLABβ—†
Pandasβ—†
NumPyβ—†
PyTorchβ—†
TensorFlowβ—†
Scikit-learnβ—†
Machine Learningβ—†
Deep Learningβ—†
Multi-Agent Systemsβ—†
NLPβ—†
LLMsβ—†
GPT-4oβ—†
RAGβ—†
XGBoostβ—†
GenAIβ—†
A/B Testingβ—†
Causal Inferenceβ—†
ETL / ELTβ—†
Apache Airflowβ—†
Apache Beamβ—†
Dataflowβ—†
Pub/Subβ—†
DBTβ—†
MLflowβ—†
CI/CDβ—†
GitHub Actionsβ—†
GCP Vertex AIβ—†
AWS SageMakerβ—†
Azureβ—†
BigQueryβ—†
Snowflakeβ—†
Dataprocβ—†
Cloud Buildβ—†
Dockerβ—†
Tableauβ—†
Power BIβ—†
Looker Studioβ—†
Data Storytellingβ—†
Neo4jβ—†
Azure AI Searchβ—†
MongoDBβ—†
Pineconeβ—†
LangChainβ—†
LlamaIndexβ—†
Vector DBsβ—†
04.

Projects

πŸ«‚

Maez β€” A Digital Companion

A new kind of AI companion that grows, learns, and bonds with a single user for life. Runs locally on consumer hardware with persistent memory, a self-evolving cognitive loop, and proposal-based autonomy β€” the user owns the credentials, Maez owns the intent.

Lifetime bond, one userRuns on-device (RTX 4090)Self-evolving cognition
Local LLMsllama.cppPersistent MemoryAgentic SystemsRAGPythonCUDASelf-Improvement
πŸ“ˆ

TrendScope

AI agent for content strategy β€” analyzes live YouTube trends, scores each by velocity, engagement, and competition, then delivers a ranked action plan with titles, hooks, and timing rationale via a ReAct reasoning loop.

ReAct agent loopCustom scoring algorithmReal-time trend signals
ReAct AgentFastAPINext.jsLiteLLMGPT-4o-miniYouTube APIPython
πŸŽ™οΈ

AI Voice FAQ Assistant

Conversational FAQ system powered by Google Gemini Pro with a full RAG pipeline β€” enabling semantic search over a product knowledge base with real-time voice interaction.

RAG pipelineSemantic searchVoice interface
Gemini ProRAGLlamaIndexPineconeFastAPINLP
πŸ›‘οΈ

Real-time Fraud Detection System

Production-grade fraud detection engine using Neo4j graph relationships and XGBoost β€” processing 5M+ daily transactions with 18% reduction in false positives.

5M+ daily txns18% fewer false positivesGraph-based detection
Neo4jXGBoostPySparkAWSGraph MLStreaming
πŸ’Š

Gym Aesthetic Trap

NLP research project using LDA topic modeling to analyze online discourse around SARMs and steroid usage β€” uncovering themes, risk perception patterns, and community sentiment from bodybuilding forums.

LDA topic modelingCommunity sentimentForum analysis
NLPLDAPythonTopic ModelingScikit-learnReddit
πŸ“

Thermal Error ML Modeling

Published Springer research on machine learning compensation strategies for thermal deformation in precision machine tools β€” achieving state-of-the-art accuracy in error prediction.

Springer publicationOct 2022Precision manufacturing
Machine LearningMATLABRegressionManufacturingSpringer
05.

Education

UMD

Master of Science – Information Systems

University of Maryland

Robert H. Smith School of Business

πŸ“ College Park, MD

GCPRequirements GatheringMLBusiness Intelligence
IIIT

Bachelor of Technology – Engineering

IIIT Chennai

Indian Institute of Information Technology

πŸ“ Chennai, India

PythonMATLABMachine LearningResearch
06.

Certifications

πŸ”·

Neo4j Graph Data Science Certification

Neo4j

Issued Feb 2025
☁️

AWS Certified AI Practitioner (AIF-C01)

Amazon Web Services

Issued 2024
πŸ“Š

BCG Data Science Job Simulation

Boston Consulting Group Γ— Forage

Issued Feb 2025
07.

Publications

Mathematical Modeling of Thermal Error Using Machine Learning

Springer Β· Oct 6, 2022

Research on thermal error modeling in machine tools using machine learning algorithms to identify the most effective compensation strategies for linear expansion and deformation caused by heat inputs from internal and external sources.

Machine LearningThermal ModelingManufacturingSpringer
08.

Model Card

πŸ€–
ramidoz/data-scientist-v4
updated jul 2026 Β· deployed to production since 2021
humanproduction-readyself-improving

// specifications

parameters
4+ years of production experience
architecture
human Γ— (ML + GenAI + product sense)
context_window
always open
training_data
healthcare Β· e-commerce Β· non-profit Β· enterprise
fine_tuned_on
GCP Vertex AI Β· AWS SageMaker
inference_hardware
coffee β˜• + RTX 4090
temperature
0.7 β€” creative but reliable
license
open_to_hire

intended use

  • βœ“Shipping LLM/RAG systems to production
  • βœ“Real-time anomaly & fraud detection
  • βœ“Graph analytics on connected data
  • βœ“MLOps: versioning, drift detection, CI/CD
  • βœ“Turning ambiguous business questions into models

out of scope

  • βœ—Dashboards nobody looks at
  • βœ—Models that never leave the notebook
  • βœ—"We'll clean the data later"
  • βœ—Meetings that should have been a Slack message

// eval results (measured in production, not on a leaderboard)

ETL latency reduction60%
Manual review time saved (RAG assistant)40%
Undetected fraud reduction (real-time)20%
False positives cut (Neo4j + XGBoost)18%
Forecast accuracy gain (Vertex AI)15%

known limitations

  • Β· may overfit to interesting problems
  • Β· inference quality degrades without coffee
  • Β· cannot resist optimizing a slow query
deploy_this_model( ) β†’
09.

Contact

Let's build something great together.

I'm currently open to Data Scientist, AI Engineer, and ML Engineer roles. Whether you're hiring, have a question about my work, or just want to talk data β€” I'd genuinely love to hear from you.

$send_message --to rohit --priority high

transmission routed via your default mail client

Usually replies within a dayΒ·rohitananthan123@gmail.com