Precision agriculture meets predictive intelligence — forecast crop suitability and yield from soil, climate, and nutrient parameters.
Topics: precision-agriculture · agricultural-ai · classification · data-science · deep-learning · machine-learning · neural-networks · random-forest-regressor · streamlit · crop-yield-forecasting
This application brings machine learning directly into the hands of farmers, agronomists, and agricultural researchers. By feeding a trained ensemble model with soil composition data, atmospheric conditions, and nutrient profiles, the system returns a ranked recommendation of the most suitable crops for the given environment — along with confidence scores and feature-level explanations that make the reasoning transparent and actionable.
The interface is built entirely in Streamlit, allowing zero-configuration deployment on any machine or cloud platform. The prediction pipeline is backed by a RandomForest or XGBoost model trained on the Crop Recommendation Dataset, with preprocessing handled by a scikit-learn Pipeline object that ensures consistent feature scaling between training and inference. A CSV batch mode allows agronomists to analyse entire fields in one operation rather than row by row.
Beyond prediction, the application includes an EDA module with interactive visualisations of the training dataset — heatmaps of feature correlations, box plots of nutrient distributions by crop type, and scatter plots of temperature vs rainfall coloured by crop category — giving users the context to interpret model outputs with domain knowledge.
Crop selection is one of the highest-stakes decisions in agriculture. A mismatch between crop requirements and soil or climatic conditions leads to yield loss, water waste, and financial hardship — particularly for smallholder farmers without access to agronomic advisory services. This project was built to democratise access to data-driven crop selection, putting a well-trained predictive model into a form that requires no technical expertise to use.
The system follows a classic ML inference pipeline:
User Input (N, P, K, pH, temp, humidity, rainfall)
│
Preprocessing Pipeline (StandardScaler / MinMaxScaler)
│
Trained Classifier (RandomForest / XGBoost / SVM)
│
Top-3 Crop Recommendations + Confidence Scores
│
Feature Importance Chart (SHAP / model.feature_importances_)
Model training is done offline in a Jupyter notebook; the serialised model (joblib) is loaded at app startup. All input validation is handled by Streamlit's built-in widget constraints before any inference is triggered.
Sliders and number inputs for Nitrogen (N), Phosphorus (P), Potassium (K), soil pH, temperature, humidity, and rainfall — all with validated ranges derived from the training data distribution.
A trained RandomForest or XGBoost model outputs a ranked list of crop recommendations with per-class probability scores, giving users confidence in the top suggestion.
Each prediction is accompanied by a SHAP waterfall chart showing which input features most strongly influenced the recommendation — making the model interpretable for non-technical users.
An embedded exploratory data analysis tab with correlation heatmaps, feature distribution plots, and crop-wise box plots for soil and climate parameters.
Upload a CSV file with multiple field samples; the app processes all rows through the same pipeline and returns a downloadable results table.
Switch between trained models (RF, XGBoost, SVM, Logistic Regression) and compare their accuracy, precision, recall, and F1-score on the held-out test set.
Optional season input (Kharif / Rabi / Zaid) adjusts feature weighting to account for seasonal temperature and rainfall norms.
No external API calls required at inference time — all model weights are bundled and loaded locally, ensuring the app works without internet access.
| Library / Tool | Role | Why This Choice |
|---|---|---|
| Streamlit | Web application framework | Zero-config Python-native UI with widget state management |
| scikit-learn | ML model training and pipeline | Pipeline, RandomForestClassifier, preprocessing transformers |
| XGBoost | Gradient-boosted tree classifier | High-accuracy alternative to RandomForest on tabular data |
| SHAP | Model explainability | TreeExplainer for feature importance visualisation |
| pandas | Data manipulation | CSV loading, filtering, batch processing |
| NumPy | Numerical operations | Array operations for feature vectors |
| Plotly | Interactive visualisation | Bar charts, scatter plots, heatmaps with zoom/hover |
| joblib | Model serialisation | Fast pickle-based model save/load |
Key packages detected in this repo:
streamlit·pandas·numpy·scikit-learn·joblib·matplotlib·seaborn·xgboost·lightgbm·tensorflow
- Python 3.9+ (or Node.js 18+ for TypeScript/JS projects)
pipornpmpackage manager- Relevant API keys (see Configuration section)
# 1. Clone and enter the repository
git clone https://github.com/Devanik21/ADV-crop-prediction-app.git
cd ADV-crop-prediction-app
# 2. Create and activate a virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# 3. Install dependencies
pip install streamlit scikit-learn xgboost shap pandas numpy plotly joblib
# 4. Launch the application
streamlit run app.py# Run the app
streamlit run app.py
# Batch prediction from CLI (if batch_predict.py is present)
python batch_predict.py --input field_samples.csv --output results.csv
# Retrain the model with new data
python train_model.py --data crop_data.csv --model rf --save model.pkl| Variable | Default | Description |
|---|---|---|
MODEL_PATH |
model.pkl |
Path to the serialised classifier |
SCALER_PATH |
scaler.pkl |
Path to the fitted preprocessing scaler |
TOP_K_CROPS |
3 |
Number of top recommendations to display |
SHAP_ENABLED |
True |
Toggle SHAP explanation computation |
Copy
.env.exampleto.envand populate all required values before running.
ADV-crop-prediction-app/
├── README.md
├── requirements.txt
├── Insights.py
├── about.py
├── analysis.py
├── analyze.py
├── comparison.py
├── crop prediction.ipynb
├── Crop_recommendation.csv
└── ...
- Integration with real-time soil sensor APIs (IoT-connected field probes)
- Geo-aware recommendations using satellite NDVI data for the input coordinates
- Yield quantity prediction (regression) alongside crop-type classification
- Mobile-optimised Streamlit layout for use on smartphones in the field
- Multi-language UI support (Hindi, Bengali, Tamil) for regional accessibility
Contributions, issues, and feature requests are welcome. Please:
- Fork the repository
- Create a feature branch (
git checkout -b feature/your-feature) - Commit your changes (
git commit -m 'feat: add your feature') - Push to your branch (
git push origin feature/your-feature) - Open a Pull Request
Please follow conventional commit messages and ensure any new code is documented.
The model was trained on the publicly available Crop Recommendation Dataset. Performance metrics (Accuracy ~99% on test set for RF) reflect the dataset's characteristics and may vary on out-of-distribution field data. Always validate recommendations with local agronomic expertise.
Devanik Debnath
B.Tech, Electronics & Communication Engineering
National Institute of Technology Agartala
This project is open source and available under the MIT License.
Crafted with curiosity, precision, and a belief that good software is worth building well.