Metadata-Version: 2.4
Name: wxcspider
Version: 0.1.0
Summary: 高性能JSON/JSONL数据处理库
Author-email: wxcSpider Team <wxcspider@example.com>
License: MIT
Project-URL: Homepage, https://pypi.org/project/wxcspider/
Project-URL: Documentation, https://pypi.org/project/wxcspider/
Keywords: json,jsonl,data-processing,spider,crawler,wxcspider
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: orjson>=3.9.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: flake8>=6.0.0; extra == "dev"
Requires-Dist: mypy>=1.0.0; extra == "dev"
Provides-Extra: full
Requires-Dist: pandas>=2.0.0; extra == "full"
Requires-Dist: openpyxl>=3.0.0; extra == "full"
Dynamic: license-file

# wxcspider

🚀 高性能 JSON/JSONL 数据处理库

专为爬虫数据处理设计，提供流式、链式 API，支持处理 GB 级大文件。

## ✨ 特性

- ⚡ **高性能** - 基于 orjson 的快速 JSON 解析
- 🌊 **流式处理** - 处理 GB 级文件不爆内存
- 🔗 **链式 API** - 优雅的函数式编程风格
- 🔧 **功能完整** - 过滤、转换、去重、合并、聚合一应俱全
- 📁 **批量处理** - 支持文件夹递归处理
- 📦 **易于扩展** - 插件化架构，轻松自定义处理器

## 📦 安装

```bash
pip install wxcspider
```

## 🚀 快速开始

### 基础用法

```python
from wxcspider import Pipeline

# 链式处理数据
(Pipeline()
    .read('data.jsonl')                          # 读取 JSONL 文件
    .filter(lambda x: x['price'] > 100)          # 过滤价格 > 100 的数据
    .transform(lambda x: {**x, 'price': float(x['price'])})  # 转换价格为浮点数
    .write('output.jsonl'))                      # 输出结果
```

### 文件夹批量处理

```python
from wxcspider import FolderProcessor

# 处理文件夹下所有 JSONL 文件（支持递归）
processor = FolderProcessor('data_folder', pattern='*.jsonl', recursive=True)

# 统计信息
stats = processor.statistics()
print(f"找到 {stats['files_found']} 个文件，共 {stats['total_records']} 条记录")

# 合并所有文件并去重
result = processor.merge_all('merged_output.jsonl', dedupe_key='id')
print(f"合并后: {result['total_records']} 条记录")
```

### 数据提取

```python
from wxcspider import DataExtractor

# 从 JSONL 文件提取指定字段
extractor = DataExtractor('products.jsonl')

# 提取单个字段
prices = extractor.extract_field('price')

# 提取多个字段
data = extractor.extract_fields(['id', 'name', 'price'])

# 提取去重后的唯一值
unique_categories = extractor.extract_unique('category')
```

### 字段操作

```python
from wxcspider import Pipeline

(Pipeline()
    .read('products.jsonl')
    .select(['id', 'name', 'price'])             # 选择特定字段
    .rename({'price': 'product_price'})          # 重命名字段
    .write('cleaned.jsonl'))
```

### 数据去重

```python
from wxcspider import Pipeline
from wxcspider.processors import DedupeProcessor

# 基于 ID 去重
(Pipeline()
    .read('data.jsonl')
    .filter(lambda x: x.get('id'))               # 确保有 ID 字段
    .transform(DedupeProcessor(key='id').process)
    .write('deduped.jsonl'))
```

## 🔧 核心功能

### 文件夹批量处理

- `FolderProcessor` - 文件夹批量处理器
  - **递归扫描** - 自动查找文件夹下所有 JSONL 文件
  - **批量合并** - 合并多个文件，可选去重
  - **批量统计** - 统计文件数、记录数、文件大小
  - **批量去重** - 对每个文件单独去重，保持目录结构
  - **批量过滤** - 批量过滤文件内容
  - **数据采样** - 从所有文件中采样数据
  - **按字段分割** - 按字段值分割成多个文件

### 数据读取

- **JSONL 流式读取** - 逐行读取，内存占用小
- **JSON 数组读取** - 支持标准 JSON 数组格式
- **压缩文件支持** - 自动处理 .gz 压缩文件
- **多文件合并** - 一次读取多个文件

### 数据处理

- **过滤器** - 条件过滤、正则表达式、范围过滤
- **转换器** - 字段选择、重命名、类型转换、嵌套展开
- **去重器** - 基于字段去重、自定义去重逻辑
- **合并器** - 多数据源合并、SQL 风格 Join
- **聚合器** - 分组统计、自定义聚合
- **排序器** - 单/多字段排序、Top N 查询
- **验证器** - 规则验证、Schema 验证

### 数据提取

- **字段提取** - 快速提取指定字段
- **嵌套字段提取** - 支持点号路径（如 `user.profile.name`）
- **去重提取** - 提取唯一值
- **多字段提取** - 批量提取多个字段

## 📚 更多示例

查看 [examples](examples/) 目录获取更多使用示例：

- `basic_usage.py` - 基础用法演示
- `advanced_usage.py` - 高级功能演示
- `extract_demo.py` - 数据提取演示

## ⚡ 性能优化建议

1. **使用流式处理** - 避免使用 `collect()` 加载全部数据到内存
2. **合理设置批次大小** - 根据数据大小调整 `batch()` 参数
3. **优先使用内置处理器** - 比自定义 lambda 函数更快

## 🔌 扩展开发

### 自定义处理器

```python
from wxcspider.processors import BaseProcessor

class MyProcessor(BaseProcessor):
    def process(self, item: dict) -> dict:
        # 自定义处理逻辑
        item['processed'] = True
        return item

# 使用
(Pipeline()
    .read('data.jsonl')
    .transform(MyProcessor().process)
    .write('output.jsonl'))
```

## 📋 依赖

- Python >= 3.8
- orjson >= 3.9.0 (高性能 JSON 库)

可选依赖：
```bash
pip install wxcspider[full]  # 安装所有可选依赖
```

- pandas >= 2.0.0 (用于高级数据分析)
- openpyxl >= 3.0.0 (用于 Excel 支持)

## 📄 许可证

MIT License

## 🤝 贡献

欢迎提交 Issue 和 Pull Request！

## 📝 更新日志

### v0.1.0 (2026-08-23)

- 初始版本发布
- 支持基础的读取、过滤、转换、去重、合并功能
- 流式处理支持
- 文件夹批量处理功能
- 数据提取工具
