Metadata-Version: 2.4
Name: amazon_product_search_v2
Version: 0.1.2
Summary: A library to search products on Amazon without using the PA API
Home-page: https://github.com/ManojPanda3/amazon-product-search
Author: Manojpanda
Author-email: Manojpanda <manojpandawork@gmail.com>
License: Copyright (c) 2025 Manoj Panda
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/ManojPanda3/amazon-product-search
Keywords: amazon,product,search,scraping,web scraping
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.7
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.7
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: beautifulsoup4>=4.11.0
Requires-Dist: requests>=2.28.0
Requires-Dist: lxml>=4.9.0
Dynamic: author
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# 🛍️ Amazon Product Search Library 📦

## Overview

Tired of manually browsing Amazon for the best deals? 🌐 Meet **Amazon Product Search** — your trusty Python library to scrape product details from Amazon's search results with just a few lines of code. Powered by **BeautifulSoup4 (bs4)**, **Requests**, and **multithreading** for speed, this library helps you efficiently gather product titles, prices, reviews, images, and direct links. 🎉

### Key Features

- **Product Search:** Search for products by name, type, brand, and price range. 📱💻
- **Detailed Data:** Scrape titles, prices (+currency), reviews (+count), images, and URLs. 🎯
- **Pagination:** Navigate Amazon pages via `page` param and get `current_page` / `total_pages` from `data-csa-c-content-id="pagination-button"`. 📄
- **Fast and Efficient:** Session reuse (keep-alive), `SoupStrainer` partial parsing, and `ThreadPoolExecutor` for extraction.
- **Easy-to-use:** Simple API + context-manager + backward-compatible iteration. ✨
- **Lightweight & Compatible:** Python 3.7–3.14, relaxed deps (`beautifulsoup4>=4.11`, `requests>=2.28`, `lxml>=4.9`). No heavy frameworks.

## Setup 🛠️

### 1. Install via PyPI (Recommended) 🧑‍💻

```bash
pip install amazon-product-search-v2
```

Installs the latest stable release. **Package name on PyPI is `amazon-product-search-v2`**.

### 2. Install via GitHub (For Developers) 🦸‍♂️

```bash
git clone --depth 1 https://github.com/ManojPanda3/amazon-product-search
cd amazon-product-search
pip install -e .
# or with venv: python -m pip install -e .
```

## Usage 📚

### Import

```python
from amazon_product_search import Amazon, AmazonProduct, AmazonResult
# Package is amazon-product-search-v2 on PyPI, but import stays amazon_product_search
```

### Basic Search

```python
from amazon_product_search import Amazon

# Session is reused across searches (keep-alive). Use as context manager to auto-close.
with Amazon() as amazon:
    result = amazon.search("thinkpad", productType="electronics", page=1)
    # or: amazon = Amazon(); result = amazon.search(...); amazon.close()

print(result.current_page, "/", result.total_pages)
for prod in result.products:
    print(prod.title, prod.price, prod.currency)
    print(prod.get())  # dict with title/link/review/review_numbers/currency/price/image
```

Or one-off:

```python
amazon = Amazon(is_debuging=False, workers=4)
result = amazon.search(productName="iPhone", productType="electronics", brand="Apple", priceRange="80000-100000", page=2)
# result is AmazonResult — also iterable/len-compatible
print(len(result))          # == len(result.products)
for product in result:      # iterates products directly
    print(product.get())
```

### Parameters for `search()`

- `productName` (str, **required**): Search term (e.g. `"thinkpad"`, `"laptop"`).
- `productType` (str, optional): `i` filter (e.g. `"electronics"`, `"books"`).
- `brand` (str, optional): Brand filter.
- `priceRange` (str, optional): `"min-max"` (e.g. `"100-200"`).
- `page` (int, optional, default `0`): **New in v0.1.2** — 1-indexed page. `0` or `1` = first page (no `page` param sent), `2` → `?page=2`.

### Returns

`AmazonResult` dataclass:

```python
@dataclass
class AmazonResult:
    products: list[AmazonProduct]
    current_page: int
    total_pages: int
    def get() -> dict: ...  # {products: [...], current_page, total_pages}
```

Each `AmazonProduct`:

```python
@dataclass
class AmazonProduct:
    title: str | None
    link: str | None          # https://www.amazon.com/dp/...
    review: str | None        # "4.5 out of 5 stars"
    review_numbers: str | None # "1234" (parentheses stripped)
    price: str | None         # "12.99"
    currency: str | None      # "$" / "₹" etc. (split on \u00a0)
    image: str | None
    def get() -> dict: ...
```

Backward compat: `for p in result`, `len(result)`, `result[i]` all proxy to `result.products`.

### Examples

#### Paginated search (new syntax)

```python
from amazon_product_search import Amazon
import json

with Amazon() as amazon:
    # Page 1
    r1 = amazon.search("thinkpad", productType="electronics", page=1)
    print(f"Page {r1.current_page} of {r1.total_pages} — {len(r1)} products")
    # Page 2
    r2 = amazon.search("thinkpad", productType="electronics", page=2)
    print(json.dumps(r2.get(), indent=2))
```

#### Iterate all pages

```python
amazon = Amazon()
page = 1
all_products = []
while True:
    res = amazon.search("laptop", page=page)
    all_products.extend(res.products)
    if res.current_page >= res.total_pages:
        break
    page += 1
print(f"Collected {len(all_products)} across {res.total_pages} pages")
amazon.close()
```

#### Old-style loop (still works)

```python
res = amazon.search("iPhone")
for product in res:  # or res.products
    d = product.get()
    print(f"Title: {d['title']}")
    print(f"Price: {d['currency']}{d['price']}")
    print(f"Review: {d['review']} ({d['review_numbers']})")
    print(f"Link: {d['link']}")
    print("-" * 40)
```

## How It Works 🔍

1. **URL building:** `urllib.parse.urlencode` safely encodes `k`, `i`, `brand`, `price`, `page`.
2. **Request:** `requests.Session` reuse (keep-alive, header persistence), 10s timeout, `RequestException` handling, `networkidle` not needed.
3. **Parse products:** `SoupStrainer("div", {"data-component-type":"s-search-result"})` + `lxml` — only product divs are parsed.
4. **Parse pagination:** `SoupStrainer("div", {"data-csa-c-content-id":"pagination-button"})` → reads `span.s-pagination-selected` (current) and max `a/span.s-pagination-item` numeric (total, e.g. `260`).
5. **Extract:** `ThreadPoolExecutor(max_workers=4)` concurrently runs `__extract_data` (title/link/review/price/image).
6. **Return:** `AmazonResult(products, current_page, total_pages)`.

## What's New in v0.1.2

- **Pagination:** `Amazon.search(..., page=N)` + `AmazonResult.current_page / total_pages` via `data-csa-c-content-id="pagination-button"`.
- **Performance:** `requests.Session` keep-alive, `SoupStrainer` partial parsing, early-exit on empty results, cached `find()` in `__get_title`.
- **Robustness:** Broader `RequestException` catch, NBSP-safe price split, `image.get("src")` type-safe, `review_numbers` paren-stripping.
- **Compatibility & Lightweight:** `from __future__ import annotations` for Python 3.7–3.14; deps relaxed to `beautifulsoup4>=4.11`, `requests>=2.28`, `lxml>=4.9` (was pinned `==`).
- **DX:** `AmazonResult` iterable/len/indexable, `.get()` dict, `Amazon` context manager (`with Amazon() as a:` + `close()`), version bump 0.1.1→0.1.2.

## Important Notes ⚠️

- **Rate Limiting:** Amazon may block frequent requests. Add delays / proxies for bulk scraping.
- **ToS:** Scraping may violate Amazon ToS — personal/educational use only.
- **Site Changes:** Selectors (`data-cy="title-recipe"`, `data-component-type="s-product-image"`, etc.) may need updates if Amazon changes markup.

## Troubleshooting 🛠️

1. `ValueError: Error product Name is required` — provide `productName`.
2. `ValueError: Error while geting data from Amazon` — network/blocked; try `Amazon(is_debuging=True)` for logs.
3. Empty `products` — no match or HTML changed; check pagination (`total_pages`) and try different `page`.
4. `None` fields — normal (Amazon varies per product).
5. `ModuleNotFoundError: No module named 'amazon_product_search'` — `pip install amazon-product-search-v2` in correct venv.

## Contributing 🤝

PRs welcome! Open an issue or submit a pull request.

## License 📜

MIT — see [LICENSE](LICENSE).
