Turning real-world datasets into production-ready AI solutions. Hands-on experience in ML, deep learning, and full-stack deployment.
Data Scientist with AI/ML internship experience and full-stack web development background. Proficient in Python, SQL, and scikit-learn. Skilled at deploying predictive apps using Flask and Streamlit.
Passionate about Generative AI, RAG, and Large Language Models.
My software engineering background gives me a distinct advantage. I don't just build ML models — I deploy and integrate them into production-ready applications.
Python, SQL
Scikit-learn, XGBoost, TensorFlow, Keras, PyTorch, CNN, LSTM, Transformers
Pandas, NumPy, Matplotlib, Seaborn
LangChain, LlamaIndex, RAG, Hugging Face, OpenAI API, Vector Embeddings
Flask, Streamlit, Node.js, REST APIs, MySQL, PostgreSQL, MongoDB
Docker, Vercel, Render, Git, GitHub

Classifies sonar signals as rock or mine using the UCI Connectionist Bench dataset. A Logistic Regression model processes 60 frequency-band features. A Flask REST API serves predictions to a custom instrument-panel frontend. A live Canvas strip-chart renders the 60-band waveform in real time.
Estimates rainfall (mm) from live atmospheric readings. Users adjust six meteorological inputs — temperature, dew point, humidity, sea-level pressure, visibility, and wind. The app returns a real-time prediction with an uncertainty range and severity category. No more static outputs buried in a Jupyter cell.

Predicts telecom subscriber churn using the IBM Telco dataset (7,043 customers, 19 features). Three classifiers were compared via 5-fold cross-validation on SMOTE-balanced data. Random Forest achieved ~78% test accuracy. A custom "diagnostic console" frontend features an animated churn-risk gauge and feature-importance breakdown. Non-technical retention teams can score subscribers and understand why the model flagged them.
Python, pandas, scikit-learn (Random Forest, Decision Tree), XGBoost, imbalanced-learn (SMOTE)
Flask, Flask-CORS, REST API (/api/predict, /api/schema, /api/health)
HTML5, CSS3, vanilla JavaScript, SVG-based animated gauge
Vercel (serverless), pickle serialization, pinned dependencies
SVM classifier on the PIMA Indians Diabetes dataset (768 records, 8 features). Achieved 76.6% test accuracy and ROC-AUC = 0.82. Flask REST API with custom real-time visualization frontend. Deployed on Vercel.
Tools: Python, scikit-learn, pandas, NumPy, Flask, Flask-CORS, HTML/CSS/JS, Vercel
Gradient Boosting on 8,000 listings across 30 locations. R² = 0.97, MAE = 5.69 Lakhs. Outlier rules improved R² from 0.81 → 0.97. Benchmarked 5 models via 5-fold CV. Dockerized Flask API on Render.
Tools: Python, Pandas, NumPy, Scikit-Learn, Flask, Gunicorn, Docker, Render
July 2026 – Current
May 2022 – Jan 2024
College of Engineering & Technology, Bhubaneswar | 2018–2021 | CGPA: 8.2
Central Tool Room & Training Center, Bhubaneswar | 2015–2018 | 65%
Real Estate Price Prediction
Diabetes Prediction (SVM)
Diabetes Prediction Model
Customer Churn (Random Forest)
Open to data science, ML engineering, and AI roles. Available for remote and on-site opportunities.
+91 7978467423
linkedin.com/in/sibashispatnaik2000
github.com/Sibashis216
"I don't just build models — I build solutions that deliver measurable value."
Explore my projects, review my code, and let's discuss how I can contribute to your data team.
© 2025 Sibashis Patnaik · Data Scientist · Berhampur, Odisha, India
Sibashis Patnaik — Data Scientist