Available for Opportunities

Ameer Sohail
Shaik

Machine Learning Engineer & Data Scientist crafting intelligent systems that turn complex data into real-world business impact.

Scroll
Ameer Sohail Shaik
1
Years Experience
5+
Projects

Building the future
with data & AI

I'm a Data Science graduate student at the University of Maryland, College Park with hands-on experience building production ML systems. My work spans stock price forecasting with machine learning, and GenAI-powered applications.

I thrive at the intersection of research and engineering — taking models from Jupyter notebooks to production APIs that serve thousands of daily requests.

Location College Park, MD
University UMD - College Park
Degree MS in Data Science
Focus ML / NLP / GenAI

Work Experience

Machine Learning Engineer
Proceedit — Remote
Aug 2024 — Aug 2025
Barcelona, Spain
  • Identified that lagging sector-ETF returns outperformed 8 other candidate features (rolling volatility, RSI, MACD) in predicting next-week price direction across 200+ tickers; communicated feature importance rankings to the trading signals team, directly informing their shift from an ARIMA baseline to a supervised learning approach.
  • Designed an offline backtesting evaluation framework comparing 10 model architectures on directional accuracy, Sharpe ratio, and max drawdown, surfaced that gradient-boosted models matched deep learning precision (65%) at 12x lower inference cost, preventing over-engineering of a production LSTM pipeline.
  • Cut data preparation workflows by 92% (60+ min to 3 min per sync) by building an automated ETL pipeline integrating PostgreSQL with Google Sheets API; wrote 50+ SQL queries (joins, window functions, CTEs) to reconcile daily OHLCV data across 4 source tables, freeing 8+ analyst-hours per week.
Python LSTM XGBoost GraphQL Flask PostgreSQL CI/CD
Data Science Intern
SBP Consulting Pvt Ltd
Jan 2024 — May 2024
Hyderabad, India
  • Analyzed 12 client accounts' SAP S/4HANA records and identified recency-weighted purchase frequency and seasonal spend ratios as top revenue predictors, engineered these into an XGBoost model achieving 87% accuracy (MAPE: 13%), with results directly shaping quarterly C-suite planning recommendations.
  • Designed 5 interactive Power BI dashboards mapping sales trends, churn risk segments, and 7+ KPIs to executive questions; replaced a 2-day manual Excel reporting cycle with 4-hour automated refresh, enabling stakeholder self-service and reducing ad-hoc data requests by ~60%.
  • Reduced data preparation time by 45% by building Python scripts (pandas, NumPy) that consolidated 4 SAP modules into analysis-ready formats via standardized cleaning logic (null imputation, duplicate detection, schema validation), eliminating recurring quality errors across 3 downstream pipelines.
Scikit-learn XGBoost SQL SAP S/4HANA Power BI Pandas

Featured Projects

🏥

Diabetes Prediction using Health Indicators

Binary classification system predicting diabetes risk from 21 health indicators using ensemble methods. Engineered 18 predictive features with domain-driven feature engineering and implemented SHAP for model interpretability.

95.9% Recall
0.993 ROC-AUC
150K+ Records
PythonXGBoostRandom ForestSHAPPandas
🤖

AutoML-ify

End-to-end AutoML web application enabling non-technical users to build ML models without coding. Supports 5+ algorithms with automated hyperparameter tuning and comprehensive model evaluation dashboards.

5+ Algorithms
100K Max Rows
80% Time Saved
PythonStreamlitScikit-learnGridSearchCV
💬

Dynamic SQL Assistant

Text-to-SQL analytics engine powered by LLM that translates natural language questions into executable SQL queries. Features LangChain-based prompt orchestration with zero hallucination errors and a Pandas ETL pipeline processing 100K+ rows.

90%+ Accuracy
100% SQL Validity
<2s Processing
LangChainGroqLlama 3.3StreamlitSQLite
📊

StreamTest: A/B Testing & Causal Inference on Streaming Data

Operationalized the product question "Does genre diversification improve user satisfaction?" on 162K MovieLens users (25M ratings), running a 6-test hypothesis suite (Welch's t, Mann-Whitney, chi-square, ANOVA, bootstrap CIs). Designed sample-size simulations showing 8,200 users/group needed at 80% power on heavy-tailed data. Estimated the causal effect of genre diversity on ratings (+0.12 stars, 95% CI: [0.06, 0.18]) via propensity score matching controlling for 3 confounders; validated with Bayesian A/B testing (98.3% P(B>A)).

162K Users
25M Ratings
98.3% P(B>A)
PythonSciPyStatsmodelsPropensity Score MatchingBayesian A/B

SparkFlow: Scalable ETL & Feature Store Pipeline

Processed 2M+ clickstream events and 500K+ transactions through a 4-stage PySpark pipeline in under 3 minutes. Wrote sessionization via SQL window functions, built 35 aggregated features, and orchestrated data quality checks (null rates, freshness, drift detection) via an Airflow DAG with SLA alerting.

2M+ Events
35 Features Built
<3min Pipeline Run
PySparkSQLApache AirflowETLFeature Store

Technical Skills

⌨️
Languages & Databases
Python SQL Java PostgreSQL MySQL SQLite MongoDB
🧠
Data Science & ML
Pandas NumPy Scikit-learn TensorFlow XGBoost LSTM Feature Engineering EDA ETL
NLP & GenAI
LangChain Llama RAG
🚀
Tools & Deployment
Flask GraphQL Git CI/CD Docker Airflow Spark Streamlit
📊
Visualization & BI
Tableau Power BI Matplotlib Plotly Seaborn
☁️
Cloud Platforms
AWS S3 AWS Redshift AWS Lambda AWS QuickSight

My Education

2025 — Present
MS in Data Science
University of Maryland — College Park
College Park, MD
2020 — 2024
B.Tech in Computer Science & Engineering
Vellore Institute of Technology
Andhra Pradesh, India

Get in Touch

Open to Summer 2026 Data Science & ML internship opportunities. Let's build something impactful together.