Terminal-Based Data Science Workflow

· 2 min read · Syed Omar Faruk Towaha

The terminal is a powerful environment for data science work. This guide shows how to build an efficient workflow using command-line tools and Python.

Essential Tools

Quick Data Exploration

Use csvkit to quickly analyze CSV files:

# Get column names
csvcut -n data.csv

# Get basic statistics
csvstat data.csv

# Filter and query
csvgrep -c column_name -m value data.csv | csvlook

# Convert to JSON
csv2json data.csv > data.json

IPython for Interactive Analysis

IPython provides a rich interactive environment:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

# Load data
df = pd.read_csv('data.csv')

# Quick exploration
df.head()
df.describe()
df.info()

# Visualization
sns.set_style('darkgrid')
plt.figure(figsize=(10, 6))
sns.histplot(df['column'])
plt.savefig('output.png')

# Use magic commands
%timeit df.groupby('category').mean()
%matplotlib inline

Scripting with Python

Create reusable analysis scripts:

#!/usr/bin/env python3
import sys
import pandas as pd
import argparse

def analyze_data(input_file, output_file):
    """Perform data analysis and save results."""
    df = pd.read_csv(input_file)

    # Analysis
    results = df.groupby('category').agg({
        'value': ['mean', 'std', 'count']
    })

    # Save results
    results.to_csv(output_file)
    print(f"Results saved to {output_file}")

if __name__ == '__main__':
    parser = argparse.ArgumentParser()
    parser.add_argument('input', help='Input CSV file')
    parser.add_argument('-o', '--output', default='results.csv')
    args = parser.parse_args()

    analyze_data(args.input, args.output)

Automation with Make

Use Makefiles to automate workflows:

# Makefile for data pipeline
.PHONY: all clean

all: results.csv plot.png

data/raw.csv:
    curl -o $@ https://example.com/data.csv

data/clean.csv: data/raw.csv scripts/clean.py
    python scripts/clean.py $< $@

results.csv: data/clean.csv scripts/analyze.py
    python scripts/analyze.py $< -o $@

plot.png: results.csv scripts/visualize.py
    python scripts/visualize.py $< $@

clean:
    rm -f data/clean.csv results.csv plot.png

Version Control for Data

Use DVC (Data Version Control) alongside Git:

# Initialize DVC
dvc init

# Track data files
dvc add data/large_dataset.csv
git add data/large_dataset.csv.dvc

# Define pipeline
dvc run -n preprocess \
  -d data/raw.csv \
  -o data/clean.csv \
  python scripts/preprocess.py

# Reproduce pipeline
dvc repro

Conclusion

A terminal-based workflow offers speed, reproducibility, and automation. While notebooks have their place, mastering terminal tools makes you a more efficient data scientist.

// related

// prefer the terminal?

Open the terminal blog and type read terminal-workflow.