Palak Tanwar

// Source

I turn messy data into systems that are auditable, not a black box.

Computer Engineer and Data Analytics Engineering MS candidate at Northeastern University (GPA 3.94/4.0). I build backend services, REST APIs, and data pipelines, with a specialization in Machine Learning & GenAI systems.

// Profile

About

I'm a Data Analytics Engineering graduate student at Northeastern University with a Computer Engineering background from Thapar Institute of Engineering & Technology. I build software end to end — REST APIs, data pipelines, and ML/GenAI systems — grounded in a strong CS foundation of OOP, data structures, and algorithms.

My experience as a Data Analyst & Strategist at Chaitanya Charitable Trust involved owning data end to end for community programs: pulling and joining data with SQL, cleaning messy survey inputs, and turning them into reporting dashboards leadership actually relied on. Today I'm focused on backend engineering, MLOps, and applying LLMs to production systems.

Outside of coursework, I'm building an end-to-end MLOps platform, exploring NLP architectures from first principles, and looking for full-time roles starting January 2027.

Palak Tanwar Portrait

Palak Tanwar — Boston, MA

// Validate

Work Experience

One role, in depth — the data programs I owned end to end.

Jun 2020 – Jul 2024

Data Analyst & Strategist

Chaitanya Charitable Trust — Gujarat, India

  • Owned data end to end for community programs: pulled and joined data with SQL, cleaned messy survey inputs in Python and spreadsheets, and turned them into trusted reporting dashboards leadership relied on.
  • Translated vague, on-the-ground needs from 1,000+ stakeholders into clear data definitions and metrics, then used the analysis to decide which campaigns to run and where to focus, growing audience reach 35%.
  • Drove cross-functional buy-in by explaining the reasoning behind the numbers in plain language, lifting program participation 30%.

// Schema

Technical Skills

The stack behind the systems — languages, foundations, infrastructure, and ML/GenAI tooling.

Languages

Python C C++ SQL (MySQL, PostgreSQL) R NoSQL (MongoDB, Neo4j)

CS Foundations

Object-Oriented Design Data Structures & Algorithms REST Client-Server Architecture Data Modeling & Validation

Backend & Cloud

FastAPI MySQL PostgreSQL MongoDB Neo4j Docker Git Unix/Linux GCP (BigQuery, GCS) Airflow MLflow

ML / GenAI & Viz

LLMs (Google Gemini) NLP Prompt Engineering NL to SQL PyTorch TensorFlow Scikit-learn XGBoost GloVe / Embeddings SHAP SMOTE Power BI QuickSight Tableau Grafana

// Output

Featured Projects

A selection spanning NLP from first principles, production MLOps, and applied LLM systems.

★ 1st Place — 6th MLOps Expo @ Google Cambridge

Foresight ML — Corporate Financial Distress Early Warning System

Apr 2026

XGBoost · LSTM/GRU · BigQuery · SHAP · MLOps

  • Awarded 1st out of 30 teams, judged on system clarity, production-readiness, deployment pipelines, and engineering depth.
  • Built an end-to-end pipeline flagging companies at risk of financial distress 6–12 months ahead, fusing SEC filings (10-K/10-Q) with FRED macro indicators for 10,000+ U.S. public companies.
  • Engineered 45+ features in BigQuery and pandas; trained XGBoost and LSTM/GRU models tuned for high recall so few at-risk firms slip through.
  • Owned the end-to-end deployment layer — model versioning, inference schema, manifest-based lineage — plus SHAP explainability and cross-sector bias checks, keeping the risk score auditable, not a black box.
Source

LLM-Powered MLOps Platform & NL Query Assistant

Dec 2025

FastAPI · Google Gemini · MySQL · Neo4j · Streamlit

  • Built a FastAPI layer wrapping a Google Gemini natural-language-to-SQL assistant behind a multistage validation pipeline that enforces read-only access and blocks unsafe queries.
  • Designed a hybrid MySQL + Neo4j backend tracking datasets, experiments, Git commit hashes, and model artifacts for fully reproducible lineage.
  • Shipped a Streamlit dashboard that uses LLMs to auto-generate plain-English summaries of model performance and KPIs for cross-functional teams.
Source

Sentiment Classification with Classical & Neural NLP Models

Jun 2026

Python · PyTorch · GloVe · CNN · NLP

  • Implemented perceptron and logistic-regression text classifiers from scratch (no ML libraries), improving test accuracy from 76% to 85% via bigram and custom feature engineering.
  • Built a Deep Averaging Network with pretrained GloVe embeddings and a CNN with multi-width convolution filters, reaching 90% test accuracy, diagnosing overfitting from loss curves.
  • Re-engineered training to mini-batch on GPU for a 20x speedup with no accuracy loss.
Source

Patient Cohort & Readmission Analytics — MIMIC-III

Feb 2025

PCA · t-SNE · K-Means · Clustering

MIMIC-III Patient Cohort Analytics Screenshot
  • Worked with real ICU patient records under strict privacy handling to segment patients and predict hospital readmissions, an operational metric that drives healthcare cost and care decisions.
  • Reduced high-dimensional clinical data with PCA and t-SNE, then used K-means and hierarchical clustering to surface distinct patient cohorts.
Source

// Sink

Get In Touch

Seeking full-time opportunities starting January 2027. Reach out if you have a role, a project, or just want to connect.