GitHub - PetrosIbrah/Data-Retrieval: Web Scraping articles and creating Search Engines

🔍 Data Retrieval

A Python and Jupyter Notebook project that implements a full data retrieval pipeline, from web scraping raw wikipedia articles to building and evaluating multiple search engine models.

📋 Requirements

Install the required packages before running:

pip install nltk scikit-learn rank-bm25 import-ipynb

⚠️ First time only — before running Part2.ipynb, uncomment the following line in the notebook, run it once, then comment it out again:
nltk.download('stopwords')

🚀 How to Run

Clone the repository

 git clone https://github.com/PetrosIbrah/Data-Retrieval.git

Open the notebooks in order and run all cells sequentially, each part depends on the output of the previous one.
```
Part1 → Part2 → Part3 → Part4a / Part4b → Part5
```

🗂️ Project structure

The project is structured as a sequential pipeline across 5 parts:

Part	Notebook	Description
1	`Part1.ipynb`	Web scraping. Ccollect articles from the web.
2	`Part2.ipynb`	Text preprocessing. Remove punctuation and stopwords.
3	`Part3.ipynb`	Build an inverted index for all remaining words
4a	`Part4a.ipynb`	Boolean Retrieval search
4b	`Part4b.ipynb`	Vector Space Model & Probabilistic search
5	`Part5.ipynb`	Evaluation. Precision, Recall, F1, MAP

Name		Name	Last commit message	Last commit date
Latest commit History 4 Commits
Data.json		Data.json
Data2.json		Data2.json
Data3.json		Data3.json
Part1.ipynb		Part1.ipynb
Part2.ipynb		Part2.ipynb
Part3.ipynb		Part3.ipynb
Part4a.ipynb		Part4a.ipynb
Part4b.ipynb		Part4b.ipynb
Part5.ipynb		Part5.ipynb
README.md		README.md

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

🔍 Data Retrieval

📋 Requirements

🚀 How to Run

🗂️ Project structure

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

🔍 Data Retrieval

📋 Requirements

🚀 How to Run

🗂️ Project structure

About

Topics

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages