Added a comprehensive CONTRIBUTING.md file to guide contributors on how to participate in the AILint project, including setup instructions, contribution guidelines, and development standards.
13 KiB
Contributing to AILint
Thanks for your interest in building the infrastructure for AI-transparent development! This guide will help you contribute effectively.
Quick Start
# Fork and clone
git clone https://github.com/evosoftie/AILint
cd AILint
# Install dependencies
npm install # for VS Code extension
pip install -r requirements.txt # for analysis engine
# Run tests
npm test
pytest
# Create a branch
git checkout -b feature/your-feature-name
Ways to Contribute
🔍 Detection Algorithms
Help improve our ability to identify AI-generated code patterns.
What we need:
- New heuristics for specific AI tools (Copilot, Claude, ChatGPT, etc.)
- Language-specific patterns (Python, C#, JavaScript, etc.)
- Adversarial testing (how to evade detection, so we can defend against it)
How to contribute:
- Document the pattern you've observed
- Implement detection logic in
/detection-engine/heuristics/ - Provide test cases (both true positives and false positives)
- Measure precision/recall on test dataset
Example:
# detection-engine/heuristics/copilot_comment_style.py
def detect_copilot_comment_pattern(code: str) -> float:
"""
Copilot tends to generate comments with specific patterns:
- Over-documentation of obvious code
- Consistent triple-slash /// style in C#
- "Helper function to..." phrasing
Returns: confidence score 0.0-1.0
"""
score = 0.0
# Check for over-documentation
if comment_to_code_ratio(code) > 0.4:
score += 0.3
# Check for characteristic phrases
if re.search(r'Helper function to|Utility method for', code):
score += 0.2
# ... more checks
return min(score, 1.0)
🔧 IDE Integration
Expand support beyond VS Code.
Platforms we need:
- Visual Studio (priority - Microsoft stack focus)
- JetBrains IDEs (IntelliJ, PyCharm, Rider)
- Vim/Neovim plugins
- Emacs integration
Requirements:
- Hook into editor events (typing, paste, file save)
- Local processing only (no code sent to servers)
- Non-blocking (can't slow down development)
- Configurable (let teams set their own thresholds)
📊 Visualization & Reporting
Make provenance data actionable.
What we need:
- Dashboard showing AI influence trends over time
- Heatmaps of codebase by AI likelihood
- Compliance report templates (SOC2, ISO 27001)
- Integration with code review tools
📐 Standards & Formats
Help define the metadata standards.
Open questions:
- What metadata format for Git notes/attributes?
- How to encode markers in different file types?
- Interoperability with other provenance tools?
- API design for external integrations?
📚 Documentation
Clear docs are critical for adoption.
Needed:
- Setup guides for different environments
- Calibration guides (tuning false positive rates)
- Architecture decision records (ADRs)
- Case studies from early adopters
Development Guidelines
Code Style
Python:
- Follow PEP 8
- Type hints required
- Docstrings for all public functions
- Black formatter (run
black .)
TypeScript/JavaScript:
- ESLint + Prettier
- Strict mode enabled
- JSDoc comments for public APIs
C#:
- Follow Microsoft coding conventions
- XML documentation comments
- StyleCop analyzer enabled
Privacy Requirements
Every PR must consider:
- What data moves off-machine? (Answer should be: nothing, or only anonymized metadata)
- Could this be used for surveillance? (Design to prevent misuse)
- Can users opt out? (For personal projects, yes)
- Is data retention clear? (Document what's stored where)
If your change violates any of these, it won't be merged.
Testing Standards
Required tests:
- Unit tests for all detection logic
- Integration tests for IDE plugins
- Performance tests (can't slow commits by >100ms)
- Adversarial tests (evasion resistance)
Test data requirements:
- Known human code (pre-2020 commits)
- Known AI code (synthetic from various models)
- Realistic hybrid code (human + AI)
Coverage:
- Core detection: >90%
- IDE integrations: >70%
- Utilities: >80%
Commit Guidelines
Format:
type(scope): brief description
Longer explanation of what and why, not how.
Reference issues: Fixes #123
AI-Assistance: [None | Copilot-autocomplete | Claude-refactor | etc.]
Types:
feat: New featurefix: Bug fixdocs: Documentation onlyperf: Performance improvementtest: Adding testsrefactor: Code change that neither fixes nor adds features
AI-Assistance tag: We practice what we preach! Declare AI help in your commits:
None- written entirely by youCopilot-autocomplete- used autocomplete suggestionsClaude-refactor- AI helped restructure codeChatGPT-debug- AI helped find/fix bugsMixed- combination of the above
This helps us dogfood our own tools and provides real-world test data.
Architecture Overview
AILint/
├── detection-engine/ # Core analysis logic
│ ├── heuristics/ # Individual detection methods
│ ├── models/ # ML models (future)
│ ├── analysis.py # Orchestration
│ └── confidence.py # Scoring system
│
├── markers/ # Embedded marker implementation
│ ├── unicode_steg.py # Zero-width character encoding
│ ├── comment_inject.py # Comment-based markers
│ └── metadata_format.py # Standard format definition
│
├── integrations/ # IDE/platform plugins
│ ├── vscode/ # VS Code extension
│ ├── visualstudio/ # Visual Studio extension
│ ├── git-hooks/ # Pre-commit hooks
│ ├── azure-devops/ # Azure DevOps plugin
│ └── github-action/ # GitHub Action
│
├── reporting/ # Dashboard & analytics
│ ├── dashboard/ # Web UI
│ ├── cli/ # Command-line reports
│ └── templates/ # Compliance templates
│
├── tests/ # Test suite
│ ├── fixtures/ # Test data
│ ├── unit/
│ ├── integration/
│ └── adversarial/
│
└── docs/ # Documentation
├── architecture/ # ADRs and design docs
├── guides/ # User guides
└── api/ # API documentation
Review Process
-
Automated checks (must pass):
- Tests pass
- Linting clean
- Coverage maintained
- Performance benchmarks acceptable
-
Human review (two approvals required):
- Code quality
- Privacy implications
- Documentation complete
- Tests adequate
-
Maintainer review:
- Architectural fit
- Standards alignment
- Roadmap coherence
Detection Heuristics - How to Add
This is the most common contribution, so here's the detailed process:
1. Document the Pattern
Create /docs/heuristics/YOUR_HEURISTIC.md:
# Heuristic: [Name]
## Pattern Description
What AI behavior are we detecting?
## Why It Works
Why does AI generate this pattern?
## False Positive Risk
What legitimate human code looks similar?
## Evasion Resistance
How easy is it to deliberately avoid this pattern?
## Test Cases
Link to test data demonstrating pattern.
2. Implement Detection
In /detection-engine/heuristics/your_heuristic.py:
from detection_engine.base import Heuristic
class YourHeuristic(Heuristic):
"""
Brief description.
False positive rate: ~X% (based on testing)
Evasion difficulty: [Low|Medium|High]
"""
def analyze(self, code: str, metadata: dict) -> float:
"""
Analyze code and return confidence score.
Args:
code: Source code to analyze
metadata: Context (language, file type, etc.)
Returns:
Confidence score 0.0-1.0
"""
# Your logic here
pass
def explain(self, code: str) -> str:
"""
Explain why this scored high.
Returns human-readable explanation for debugging.
"""
pass
3. Add Tests
In /tests/heuristics/test_your_heuristic.py:
import pytest
from detection_engine.heuristics.your_heuristic import YourHeuristic
class TestYourHeuristic:
def test_detects_ai_pattern(self):
"""Should score high on known AI code."""
heuristic = YourHeuristic()
code = load_fixture('ai_generated/example1.py')
score = heuristic.analyze(code, {'language': 'python'})
assert score > 0.7
def test_human_code_low_score(self):
"""Should score low on human code."""
heuristic = YourHeuristic()
code = load_fixture('human_written/example1.py')
score = heuristic.analyze(code, {'language': 'python'})
assert score < 0.3
def test_edge_cases(self):
"""Handle edge cases gracefully."""
heuristic = YourHeuristic()
# Empty file
assert heuristic.analyze('', {}) == 0.0
# Minified code
# Obfuscated code
# etc.
4. Benchmark Performance
python -m detection_engine.benchmark your_heuristic
# Should complete in <10ms for typical files
# Should handle files up to 10k lines
5. Submit PR
Title: feat(heuristic): Add [pattern name] detection
Include in description:
- What pattern this detects
- Measured false positive rate
- Test coverage
- Performance benchmarks
Community
Communication Channels
- GitHub Issues: Bug reports, feature requests
- GitHub Discussions: Questions, ideas, showcase
- Discord (coming soon): Real-time chat
- Monthly calls (coming soon): Community sync
Code of Conduct
We're building infrastructure for transparency and trust. Our community must model these values:
- Assume good faith: People may disagree on approaches
- Be respectful: Attack ideas, not people
- Privacy matters: Never share others' code without permission
- Collaborate openly: Solutions over credit
Full Code of Conduct: CODE_OF_CONDUCT.md
Recognition
Contributors are recognized in:
- README.md credits
- Release notes
- Annual report (coming 2025)
Significant contributions may earn:
- Commit access
- Maintainer role
- Representation at conferences
Legal
Licensing
- All code: MIT License
- All documentation: CC BY 4.0
By submitting a PR, you agree to license your contribution under these terms.
Patents
You confirm you have the right to contribute the code and are not violating any patents or other IP rights.
Privacy of Test Data
If contributing test data:
- Must be your own code, or
- Must have explicit permission, or
- Must be public domain / open source
Never contribute proprietary code as test fixtures.
Getting Help
Stuck? Here's how to get unblocked:
- Check docs:
/docsfolder and GitHub wiki - Search issues: Someone may have asked already
- Ask in Discussions: For design questions
- Open an issue: For bugs or feature requests
- Tag maintainers: If urgent (but respect their time)
Mentorship: New to open source or AI detection? Open an issue labeled good-first-issue or mentorship-wanted and we'll help you get started.
Roadmap Input
We maintain the roadmap in ROADMAP.md.
Want to influence priorities? Open an issue labeled roadmap-input explaining:
- What capability you need
- Why it matters
- Rough implementation approach
Community needs shape what we build.
Thank You
Every contribution - code, docs, testing, ideas - moves us closer to an AI-transparent future.
You're not just building a tool. You're laying the pipes while we build the motorways.
Let's build the Bronze Age together.
Questions about contributing? Open an issue labeled contributing-question and we'll update this doc.
---
# VS Code Extension Architecture
Now the technical design for the first implementation:
## Extension Structure
vscode-extension/ ├── src/ │ ├── extension.ts # Entry point │ ├── monitors/ │ │ ├── typingMonitor.ts # Track typing patterns │ │ ├── pasteMonitor.ts # Detect paste events │ │ └── fileMonitor.ts # Watch file changes │ ├── analysis/ │ │ ├── analyzer.ts # Orchestrate detection │ │ ├── localEngine.ts # Interface to detection engine │ │ └── cache.ts # Performance optimization │ ├── git/ │ │ ├── gitIntegration.ts # Git commands │ │ ├── metadata.ts # Format metadata │ │ └── hooks.ts # Pre-commit integration │ ├── ui/ │ │ ├── statusBar.ts # Show AI likelihood in status │ │ ├── decorations.ts # Highlight suspicious code │ │ ├── webview.ts # Dashboard panel │ │ └── notifications.ts # User feedback │ └── config/ │ └── settings.ts # User preferences ├── detection-engine/ # Python analysis (bundled) │ └── [Core detection logic] ├── package.json ├── tsconfig.json └── README.md