Metadata-Version: 2.4
Name: csim
Version: 3.0.2
Summary: Code Similarity (csim) is a method designed to detect similarity between source codes
Home-page: https://github.com/EdsonEddy/csim
Author: Eddy Lecoña
Author-email: crew0eddy@gmail.com
License: MIT
Project-URL: Bug Tracker, https://github.com/EdsonEddy/csim/issues
Project-URL: Documentation, https://github.com/EdsonEddy/csim/wiki
Project-URL: Source Code, https://github.com/EdsonEddy/csim
Keywords: code analysis,similarity detection,tree parser,tree edit distance,code snippets,code comparison
Platform: All
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Development Status :: 4 - Beta
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: antlr4-python3-runtime==4.13.2
Requires-Dist: zss==1.2.0
Requires-Dist: numpy==1.26.4
Requires-Dist: apted==1.0.3
Dynamic: author
Dynamic: author-email
Dynamic: classifier
Dynamic: description
Dynamic: description-content-type
Dynamic: home-page
Dynamic: keywords
Dynamic: license
Dynamic: license-file
Dynamic: platform
Dynamic: project-url
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

# Code Similarity (csim)

Code Similarity (csim) provide a module designed to detect similarities between source code files, even when obfuscation techniques have been applied. It is particularly useful for programming instructors and students who need to verify code originality.

## Key Features

- **Source Code Similarity Analysis:** Compares source code files to determine their degree of similarity.
- **Pairwise Reporting:** Generate detailed similarity reports for all file pairs.
- **File Grouping:** Cluster similar files into groups based on a configurable threshold.
- **Flexible Search Strategies:** 
  - **Exhaustive Search:** All-pairs comparison for maximum precision
- **Advanced Analysis:** Utilizes parse trees and the tree edit distance algorithm for in-depth analysis.
- **Parse Trees:** Represents the syntactic structure of source code, enabling detailed comparisons.
- **Tree Edit Distance:** Measures the similarity between different code structures.
- **Hash-Based Pruning:** Optimizes the comparison process by reducing tree size while preserving essential structure.
- **Multi-Language Support:** Supports Python 3.13, Java 20, and C++14 source code analysis.

## Technologies Used

- **Python:** The core programming language for the tool.
- **ANTLR:** A parser generator for creating parse trees from source code.
- **apted:** A library for computing the tree edit distance (default algorithm).
- **zss:** A library for calculating the tree edit distance, alternatively to apted.
- **NumPy:** Used for efficient numerical operations.

## Installation
For the installation `pip` is required, you can either clone the repository and install it locally or install it directly from PyPI.

1.  Clone the repository:
    ```sh
    git clone https://github.com/EdsonEddy/csim.git
    ```
2.  Navigate to the project directory:
    ```sh
    cd csim
    ```
3.  Install the package:
    ```sh
    pip install .
    ```

Alternatively, you can install it directly from PyPI:

```sh
pip install csim
```


### Version Compatibility
- **Python:** 3.10–3.12 (recommended 3.11)
- **ANTLR4 Python Runtime:** 4.13.2
- **zss:** 1.2.0
- **apted:** 1.0.3
- **numpy:** 1.26.4

## Quick Start

**New to csim?** Start here: [GETTING_STARTED.md](GETTING_STARTED.md)

For detailed information about search strategies, see: [docs/STRATEGIES.md](docs/STRATEGIES.md)

csim supports three main actions: **report** (for pairwise similarity analysis), **group** (for clustering similar files), and **tree**/**view** (for visualizing a file's normalized/pruned parse tree). The tool supports Python 3.13, Java 20, and C++14 source code files.

### General Command Structure
```sh
csim <action> --path <directory> [options]
```

### Action 1: `report` - Generate Similarity Report

Generates a pairwise similarity report comparing all files in a directory.

```sh
csim report --path /path/to/directory
```

**Example Output:**
```
file1.py is similar to file2.py with similarity index: 0.95
file1.py is similar to file3.py with similarity index: 0.45
file2.py is similar to file3.py with similarity index: 0.50
```

**Options:**
- `--lang, -l`: Programming language (default: `python_3_13`). Options: `python_3_13`, `java_20`, `cpp_14`
- `--talg, -ta`: Tree edit distance algorithm (default: `apted`). Options: `zss`, `apted`

**Example with options:**
```sh
csim report --path /path/to/directory --lang java_20 --talg zss
```

### Action 2: `group` - Group Files by Similarity

Groups files by similarity using a specified threshold and strategy.

```sh
csim group --path /path/to/directory --threshold 0.8
```

**Example Output:**
```
Threshold: 0.8
Total files processed: 4
Group 1 (Average Similarity: 0.98):
./file1.py
./file2.py
Group 2 (Average Similarity: 0.95):
./file3.py
./file4.py
```

#### Strategy Options

The `group` action supports two strategies for finding similar files:

##### 1. **exhaustive** (Default)
Compares every file against every other file (O(n²)). This is the most thorough approach but slower for large datasets.

```sh
csim group --path /path/to/directory --threshold 0.8 --strategy exhaustive
```

**When to use each:**
- **exhaustive**: Small datasets (< 100 files), when maximum precision is critical

#### Group Action Options

- `--threshold, -t`: Similarity threshold (0.0 to 1.0). **Required.**
- `--strategy, -s`: Grouping strategy (default: `exhaustive`). Options: `exhaustive`
- `--lang, -l`: Programming language (default: `python_3_13`). Options: `python_3_13`, `java_20`, `cpp_14`
- `--talg, -ta`: Tree edit distance algorithm (default: `apted`). Options: `zss`, `apted`

**Complete example:**
```sh
csim group --path /path/to/directory --threshold 0.9 --strategy exhaustive --lang python_3_13 --talg zss
```

### Action 3: `tree` (alias: `view`) - Visualize Parse Trees

Prints the normalized/pruned tree for a single file — the exact tree that gets passed to the tree edit distance algorithm. Useful for debugging how the normalization, collapsing, and hashing rules affect a specific file before it's compared against others.

```sh
csim tree --path /path/to/file.py --lang python_3_13
```

**Example Output:**
```
=== Normalized + Pruned Tree (input to Tree Edit Distance) ===
statements
   function_def_raw
      param [hashed:e3b0c442]
      statements
         STRING
         if_stmt
            comparison [hashed:93e10dca]
            return_stmt [hashed:337adaa9]
   assignment [hashed:118045cc]
   primary [hashed:e1b0c7ab]

Total nodes after pruning: 24
```

Rule and token names are resolved for readability, `LOOP` marks nodes collapsed under control-flow equivalence (e.g. `for`/`while`), and `[hashed:xxxxxxxx]` marks subtrees that were hashed into a single node instead of compared structurally.

**Options:**
- `--path, -p`: Path to a single source code file (**required**).
- `--lang, -l`: Programming language (default: `python_3_13`). Options: `python_3_13`, `java_20`, `cpp_14`
- `--show-raw`: Also print the raw ANTLR parse tree before normalization/pruning, for side-by-side comparison.

**Example with `--show-raw`:**
```sh
csim tree --path /path/to/file.py --lang python_3_13 --show-raw
```

### Language Support

The tool supports the following programming languages:

**Python 3.13:**
```sh
csim report --path /path/to/python/files --lang python_3_13
```

**Java 20:**
```sh
csim report --path /path/to/java/files --lang java_20
```

**C++14:**
```sh
csim report --path /path/to/cpp/files --lang cpp_14
```

### Threshold Guidance

The similarity threshold represents the structural similarity of the code (based on the Abstract Syntax Tree). Choose appropriate thresholds based on your use case:

- **0.95+**: Nearly identical code (likely plagiarism)
- **0.85-0.95**: Very similar code (probable plagiarism)
- **0.70-0.85**: Moderately similar code (review recommended)
- **<0.70**: Low similarity (likely independent work)

### Using csim as a Python Module

You can also use csim programmatically within your Python code. The library provides low-level functions for advanced use cases:

```python
from csim.utils import group_by_exhaustive_search, report_pairwise_similarity

# Example: Group files by similarity
file_names = ["file1.py", "file2.py", "file3.py"]
file_contents = [code1, code2, code3]

results = group_by_exhaustive_search(
    file_names=file_names,
    file_contents=file_contents,
    lang="python_3_13",
    threshold=0.8,
    ted_algorithm="apted"
)

print(results)
```

Or use the legacy Compare class for simple pairwise comparisons:

```python
from csim import Compare

code_a = "a = 5"
code_b = "c = 50"
similarity = Compare(name_a='example A', content_a=code_a, name_b='example B', content_b=code_b)
print(f"Similarity: {similarity}") # Output: Similarity: X.XX
```

## Documentation

- [Getting Started Guide](GETTING_STARTED.md) - Quick tutorial for new users
- [Search Strategies Guide](docs/STRATEGIES.md) - Detailed explanation of available search strategies
- [ANTLR Parser Generation](grammars/parser_gen_guide.md) - For grammar customization

## ANTLR4 Installation and Parser/Lexer Generation

This installation is not required—the generated files are already included in the project. If you'd like to review the steps to generate them yourself, see [grammars/parser_gen_guide.md](grammars/parser_gen_guide.md).

Note: The included generated files were produced by **ANTLR 4.13.2** and are compatible with the pinned runtime listed above.

## Contributing

Contributions are welcome! To contribute, please follow these steps:

1.  Fork the repository.
2.  Create a new branch (`git checkout -b feature/new-feature`).
3.  Make your changes and commit them (`git commit -am 'Add new feature'`).
4.  Push to the branch (`git push origin feature/new-feature`).
5.  Open a Pull Request.

## License

This project is licensed under the MIT License. See the [LICENSE](LICENSE) file for details.

## Support

- **Questions?** Open a [GitHub Discussion](https://github.com/EdsonEddy/csim/discussions)
- **Found a bug?** File a [GitHub Issue](https://github.com/EdsonEddy/csim/issues)
- **Want to contribute?** See [Contributing](#contributing) section

## References

For more information on the techniques and tools used in this project, refer to the following resources:

- [ANTLR](https://www.antlr.org/)
- [Parse Tree (Wikipedia)](https://en.wikipedia.org/wiki/Parse_tree)
- [Tree Edit Distance (Wikipedia)](https://en.wikipedia.org/wiki/Tree_edit_distance)
- [Locality Sensitive Hashing (Wikipedia)](https://en.wikipedia.org/wiki/Locality-sensitive_hashing)
- [MinHash (Wikipedia)](https://en.wikipedia.org/wiki/MinHash)
- [zss (PyPI)](https://pypi.org/project/zss/)
- [Hashing (Python Docs)](https://docs.python.org/3/library/hashlib.html)
- [apted (GitHub)](https://github.com/JoaoFelipe/apted)

## Third-Party Licenses

This project utilizes the following third-party libraries:

### ANTLR (ANother Tool for Language Recognition)
- **Purpose:** A parser generator used to create parse trees from source code.
- **License:** BSD 3-Clause
- **Website:** [https://www.antlr.org/](https://www.antlr.org/)
- **Repository:** [https://github.com/antlr/antlr4](https://github.com/antlr/antlr4)

### ANTLR4-parser-for-Python-3.14 by RobEin
- **Purpose:** Python 3.14 grammar for ANTLR4
- **License:** MIT License
- **Repository:** [https://github.com/RobEin/ANTLR4-parser-for-Python-3.14](https://github.com/RobEin/ANTLR4-parser-for-Python-3.14)

### zss (Zhang-Shasha)
- **Purpose:** Tree edit distance algorithm implementation for comparing tree structures
- **License:** MIT License
- **Repository:** [https://github.com/timtadh/zhang-shasha](https://github.com/timtadh/zhang-shasha)

### apted (All Path Tree Edit Distance)
- **Purpose:** Python APTED algorithm for the Tree Edit Distance, an alternative to zss
- **License:** MIT License
- **Repository:** [https://github.com/JoaoFelipe/apted](https://github.com/JoaoFelipe/apted)

