Metadata-Version: 2.4
Name: hark-cli
Version: 0.3.0
Summary: 100% offline, Whisper-powered voice notes from your terminal
Project-URL: Homepage, https://github.com/FPurchess/hark
Project-URL: Repository, https://github.com/FPurchess/hark
Project-URL: Issues, https://github.com/FPurchess/hark/issues
License: AGPL-3.0-or-later
License-File: LICENSE
Keywords: cli,faster-whisper,offline,speech-to-text,transcription,voice-notes,whisper
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: End Users/Desktop
Classifier: License :: OSI Approved :: GNU Affero General Public License v3 or later (AGPLv3+)
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows :: Windows 10
Classifier: Operating System :: Microsoft :: Windows :: Windows 11
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.11
Requires-Dist: faster-whisper>=1.0.0
Requires-Dist: librosa>=0.10.0
Requires-Dist: noisereduce>=3.0.0
Requires-Dist: numpy>=1.24.0
Requires-Dist: pulsectl>=24.0.0; sys_platform == 'linux'
Requires-Dist: pyaudiowpatch>=0.2.12.7; sys_platform == 'win32'
Requires-Dist: pyyaml>=6.0.1
Requires-Dist: sounddevice>=0.5.0
Requires-Dist: soundfile>=0.12.0
Provides-Extra: dev
Requires-Dist: pre-commit>=3.0.0; extra == 'dev'
Requires-Dist: pyrefly>=0.45.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0.0; extra == 'dev'
Requires-Dist: pytest>=7.0.0; extra == 'dev'
Requires-Dist: ruff>=0.1.0; extra == 'dev'
Requires-Dist: whisperx>=3.1.0; extra == 'dev'
Provides-Extra: diarization
Requires-Dist: whisperx>=3.1.0; extra == 'diarization'
Provides-Extra: test
Requires-Dist: pre-commit>=3.0.0; extra == 'test'
Requires-Dist: pyrefly>=0.45.0; extra == 'test'
Requires-Dist: pytest-cov>=4.0.0; extra == 'test'
Requires-Dist: pytest>=7.0.0; extra == 'test'
Requires-Dist: ruff>=0.1.0; extra == 'test'
Requires-Dist: whisperx>=3.1.0; extra == 'test'
Description-Content-Type: text/markdown

# hark 😇

[![PyPI version](https://img.shields.io/pypi/v/hark-cli)](https://pypi.org/project/hark-cli/)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue)](https://www.python.org/downloads/)
[![License: AGPL v3](https://img.shields.io/badge/License-AGPL%20v3-blue.svg)](https://www.gnu.org/licenses/agpl-3.0)
[![Code style: ruff](https://img.shields.io/badge/code%20style-ruff-000000.svg)](https://github.com/astral-sh/ruff)

> 100% offline voice notes from your terminal

### Use Cases

- **Voice-to-LLM pipelines** — `hark | llm` turns speech into AI prompts instantly
- **Meeting minutes** — Transcribe calls with speaker identification (`--diarize`)
- **System audio capture** — Record what you hear, not just what you say (`--input speaker`)
- **Private by design** — No cloud, no API keys, no data leaves your machine

## Features

- 🎙️ **Instant Recording** - One keypress to capture your thoughts
- 🔊 **Multi-Source Capture** - Record microphone, system audio, or both simultaneously
- ✨ **High-Accuracy Transcription** - State-of-the-art speech recognition for crystal-clear text
- 🗣️ **Speaker Diarization** - Automatically identify and label who said what
- 🔒 **Complete Privacy** - 100% offline processing, your audio never leaves your device
- 📄 **Flexible Output** - Export as plain text, markdown, or SRT subtitles
- 🌍 **Multilingual Support** - Transcribe in dozens of languages with automatic detection
- ⚡ **Blazing Fast** - Hardware-accelerated processing for near real-time results

## Installation

```bash
pipx install hark-cli
```

### System Dependencies

**Ubuntu/Debian:**

```bash
sudo apt install portaudio19-dev
```

**macOS:**

```bash
brew install portaudio
```

**Windows:**

No system dependencies required. Audio libraries are bundled with the Python packages.

## Quick Start

```bash
# Record and print to stdout
hark

# Save to file
hark notes.txt

# Use larger model for better accuracy
hark --model large-v3 meeting.md

# Transcribe in German
hark --lang de notes.txt

# Output as SRT subtitles
hark --format srt captions.srt

# Capture system audio (e.g., online meetings)
hark --input speaker meeting.txt

# Capture both microphone and system audio (stereo: L=mic, R=speaker)
hark --input both conversation.txt
```

## Configuration

Hark uses a YAML config file at `~/.config/hark/config.yaml`. CLI flags override config file settings.

```yaml
# ~/.config/hark/config.yaml
recording:
  sample_rate: 16000
  channels: 1 # Use 2 for --input both
  max_duration: 600
  input_source: mic # mic, speaker, or both

whisper:
  model: base # tiny, base, small, medium, large, large-v2, large-v3
  language: auto # auto, en, de, fr, es, ...
  device: auto # auto, cpu, cuda

preprocessing:
  noise_reduction:
    enabled: true
    strength: 0.5 # 0.0-1.0
  normalization:
    enabled: true
  silence_trimming:
    enabled: true

output:
  format: plain # plain, markdown, srt
  timestamps: false

diarization:
  hf_token: null # HuggingFace token (required for --diarize)
  local_speaker_name: null # Your name in stereo mode, or null for SPEAKER_00
```

## Audio Input Sources

Hark supports three input modes via `--input` or `recording.input_source`:

| Mode      | Description                                            |
| --------- | ------------------------------------------------------ |
| `mic`     | Microphone only (default)                              |
| `speaker` | System audio only (loopback capture)                   |
| `both`    | Microphone + system audio as stereo (L=mic, R=speaker) |

### System Audio Capture

System audio capture (`--input speaker` or `--input both`) works differently on each platform:

**Linux (PulseAudio/PipeWire):**

Uses monitor sources automatically. To verify your system supports it:

```bash
pactl list sources | grep -i monitor
```

You should see output like:

```
Name: alsa_output.pci-0000_00_1f.3.analog-stereo.monitor
Description: Monitor of Built-in Audio
```

**macOS:**

Requires [BlackHole](https://github.com/ExistentialAudio/BlackHole) virtual audio driver:

1. Install BlackHole:

   ```bash
   brew install blackhole-2ch
   ```

2. Open **Audio MIDI Setup** (in Applications → Utilities)

3. Click **+** → **Create Multi-Output Device**

4. Check both your speakers/headphones AND BlackHole 2ch

5. Set the Multi-Output Device as your default output in System Preferences → Sound

Now hark can capture system audio through BlackHole.

**Windows 10/11:**

Uses WASAPI loopback automatically. No setup required—just ensure your audio output device is working.

## Speaker Diarization

Identify who said what in multi-speaker recordings using [WhisperX](https://github.com/m-bain/whisperX).

### Setup

1. Install diarization dependencies:

   ```bash
   pipx inject hark-cli whisperx
   # Or with pip:
   pip install hark-cli[diarization]
   ```

2. Get a HuggingFace token (required for pyannote models):

   - Create account at https://huggingface.co
   - Accept model licenses:
     - https://huggingface.co/pyannote/segmentation-3.0
     - https://huggingface.co/pyannote/speaker-diarization-3.1
   - Create token at https://huggingface.co/settings/tokens

3. Add token to config:
   ```yaml
   # ~/.config/hark/config.yaml
   diarization:
     hf_token: "hf_xxxxxxxxxxxxx"
   ```

### Usage

The `--diarize` flag enables speaker identification. It requires `--input speaker` or `--input both`.

```bash
# Transcribe a meeting with speaker identification
hark --diarize --input speaker meeting.txt

# Specify expected number of speakers (improves accuracy)
hark --diarize --speakers 3 --input speaker meeting.md

# Skip interactive speaker naming for batch processing
hark --diarize --no-interactive --input speaker meeting.txt

# Stereo mode: separate local user from remote speakers
hark --diarize --input both conversation.md

# Combine with other options
hark --diarize --input speaker --format markdown --model large-v3 meeting.md
```

| Flag               | Description                                           |
| ------------------ | ----------------------------------------------------- |
| `--diarize`        | Enable speaker identification                         |
| `--speakers N`     | Hint for expected speaker count (improves clustering) |
| `--no-interactive` | Skip post-transcription speaker naming prompt         |

**Note:** Diarization adds processing time. For a 5-minute recording, expect ~1-2 minutes on GPU or ~5-10 minutes on CPU.

### Output Format

With diarization enabled, output includes speaker labels and timestamps:

**Plain text:**

```
[00:02] [SPEAKER_01] Hello everyone, let's get started.
[00:05] [SPEAKER_02] Thanks for joining. Let me share my screen.
```

**Markdown:**

```markdown
# Meeting Transcript

**SPEAKER_01** (00:02)
Hello everyone, let's get started.

**SPEAKER_02** (00:05)
Thanks for joining. Let me share my screen.

---

_2 speakers detected • Duration: 5:23 • Language: en (98% confidence)_
```

### Interactive Naming

After transcription, hark will prompt you to identify speakers:

```
Detected 2 speaker(s) to identify.

SPEAKER_01 said: "Hello everyone, let's get started."
Who is this? [name/skip/done]: Alice

SPEAKER_02 said: "Thanks for joining. Let me share my screen."
Who is this? [name/skip/done]: Bob
```

Use `--no-interactive` to skip this prompt.

### Known Issues

**Slow diarization?** The pyannote models may default to CPU inference. For GPU acceleration:

```bash
pip install --force-reinstall onnxruntime-gpu
```

See [WhisperX #499](https://github.com/m-bain/whisperX/issues/499) for details.

## Development

```bash
git clone https://github.com/FPurchess/hark.git
cd hark
uv sync --extra test
uv run pre-commit install
uv run pytest
```

## Contributing

Contributions are what make the open source community such an amazing place to learn, inspire, and create. Any contributions you make are **greatly appreciated**.

If you have a suggestion that would make this better, please fork the repo and create a pull request. You can also simply open an issue with the tag "enhancement".
Don't forget to give the project a star! Thanks again!

1. Fork the Project
2. Create your Feature Branch (`git checkout -b feature/AmazingFeature`)
3. Commit your Changes (`git commit -m 'Add some AmazingFeature'`)
4. Push to the Branch (`git push origin feature/AmazingFeature`)
5. Open a Pull Request

## License

Distributed under the [**AGPLv3 License**](LICENSE).

[![FOSSA Status](https://app.fossa.com/api/projects/git%2Bgithub.com%2FFPurchess%2Fhark.svg?type=large)](https://app.fossa.com/projects/git%2Bgithub.com%2FFPurchess%2Fhark?ref=badge_large)

## Acknowledgments

This project would not exist without the hard work of others, first and foremost the maintainers and contributors of the below mentioned projects:

- [faster-whisper](https://github.com/SYSTRAN/faster-whisper)
- [WhisperX](https://github.com/m-bain/whisperX)
