The New York company teaching the internet a harder question than "who wrote this?" - it asks whether anyone wrote it at all.
Copyleaks began, like a lot of durable companies, with a personal grievance. Around 2015, content on a website belonging to co-founder Yehonatan Bitton's family was copied and reused without permission. The theft cost web traffic, business and revenue. Bitton and his co-founder Alon Yamin - both software developers who had worked on text analysis and machine learning, including time in Israel's Unit 8200 intelligence corps - decided that the same pattern-recognition instincts used to sift signal from noise could be pointed at a more everyday problem: telling original writing from copied writing.
What started as a plagiarism checker became something larger, and it did so at an unusually good moment. Copyleaks spent years building infrastructure to compare text against enormous databases and to weigh not just matching words but semantic meaning and writing style. Then generative AI arrived. When ChatGPT reached classrooms and newsrooms, nearly every teacher, editor and compliance officer suddenly needed to ask a question the old tools were not designed for - not "did you copy this?" but "did a machine make this?" Copyleaks already had much of the machinery to answer it.
A plagiarism checker used to compare your words to a database. Now it has to ask a harder question: did a human write this at all?
Today the company sells a content-authenticity platform rather than a single tool. Its AI Detector identifies text produced by large language models - including ChatGPT, GPT-4, Gemini and Claude - and claims to catch paraphrased output at accuracy rates it puts above 99% across more than 30 languages. Its plagiarism checker remains the original engine. Codeleaks extends the same idea to software, flagging copied and AI-generated source code while surfacing the open-source licenses buried inside it, a growing headache for engineering teams. And in 2025 and 2026 the company pushed into visual media, launching AI image detection and an AI video detector that scans video and audio tracks together for signs of synthetic generation.
The customers fall into three broad camps. Educational institutions - from K-12 districts to universities - use Copyleaks to defend academic integrity, usually through integrations built directly into learning management systems like Canvas, Moodle, Blackboard and D2L Brightspace. Publishers and content creators use it to protect originality and prove their work is human-made. And enterprises use it to manage brand, legal and compliance risk as AI-written material seeps into contracts, documentation and code.
The problems it solves are all versions of the same anxiety. In a world where anyone can generate fluent text, functioning code or photorealistic images in seconds, how do you know what is original, what is licensed, and what came out of a model? Copyleaks positions itself as the verification layer for that question - the thing you run before you grade an essay, publish an article, ship a feature or trust an image.
What separates it from rivals is partly breadth and partly posture. Where many detectors handle English text and little else, Copyleaks leans on multilingual coverage and deep enterprise integrations. And where most tools hand back a bare confidence score, Copyleaks in 2025 launched AI Logic, a feature built to explain why a passage was flagged rather than just how likely it is to be AI. That distinction matters most in education, where a false positive is not a statistic - it is a student wrongly accused.
The company is candid, at least by industry standards, about the limits. Independent studies have found strong detection of unedited AI content but noted that human editing can blunt any detector's accuracy. Copyleaks' answer is not to claim perfection but to make its judgments legible and to keep widening the net across languages and media.
The business itself is a study in the value of unglamorous problems. Copyleaks reported revenue growth of 1,310% over three years, enough to land it at #330 on the Inc. 5000 and among Connecticut's fastest-growing private companies. It raised a $6M Series A in April 2022 led by JAL Ventures, following an earlier seed round backed by the State of Connecticut's investment fund, Connecticut Innovations. With roughly 92 employees and an estimated $5M in annual revenue, it is not the biggest name in its category - Turnitin looms large in education - but it has carved a distinct position by refusing to stay a single-purpose tool.
In a world of synthetic everything, verification stops being a feature and starts being infrastructure.
That is the bet underneath everything Copyleaks builds. The volume of machine-generated content is only going up, and it is spreading from text into code, images and video faster than institutions can adapt. Someone has to build the trust layer - the tool that quietly answers "is this real?" before decisions get made on top of it. Copyleaks is trying to be that tool, in as many languages and formats as the problem demands.
Identifies text from LLMs like ChatGPT, GPT-4, Gemini and Claude - including paraphrased content - with claimed 99%+ accuracy across 30+ languages.
The original engine. Compares writing against vast databases, weighing semantic meaning and style, and returns detailed similarity reports.
Detects AI-generated and plagiarized source code and surfaces the software licenses embedded within it - a compliance tool for engineering teams.
Explains why a passage was flagged rather than handing back a bare score. Rolled out across Canvas, Moodle, Blackboard, D2L and more.
Flags AI-generated images to support trust and transparency in visual content across education and enterprise use.
Scans video and audio tracks simultaneously to identify AI-generated and deepfake content.
Alon Yamin and Yehonatan Bitton launch Copyleaks to detect plagiarism using AI and text analysis.
Raises $1.8M led by Connecticut Innovations to grow its detection platform.
Extends detection to source code, flagging plagiarized and licensed code.
Raises $6M led by JAL Ventures to expand anti-plagiarism and AI capabilities.
Ships AI-generated text detection as ChatGPT reshapes education and publishing.
Adds transparency to detection across LMS platforms and partners with RWS on Tridion Docs.
Adds AI image and video detection - authenticity across text, code, images and video.
Series A led by JAL Ventures with Connecticut Innovations, Acadian Ventures, G2C Venture Partners, TLI Bedrock and True Blue Partners. Roughly $7-8M raised in total.
A former software developer with a background in text analysis and Israel's Unit 8200 intelligence corps, Yamin has led Copyleaks since 2015 and is the public voice on responsible AI detection.
Co-founded Copyleaks after his family's website content was stolen and reused - the copyright fight that sparked the whole company.
An engineering- and research-driven team based in New York, focused on accuracy, transparency and expanding detection into new languages and media.
It uses AI to detect plagiarism, AI-generated text, and copied or AI-generated source code, and has expanded into detecting AI-generated images and video.
It was founded in 2015 by Alon Yamin (CEO) and Yehonatan Bitton, both former software developers with backgrounds in text analysis and Israeli intelligence.
Copyleaks claims over 99% accuracy for AI-text detection across 30+ languages. Independent studies show strong results on unedited AI content but note detectors can be circumvented by human editing.
Educational institutions, publishers and content creators, and enterprises - via subscriptions, an API, and integrations with LMS platforms like Canvas, Moodle and Blackboard.
Roughly $7-8M in total, including a $6M Series A in April 2022 led by JAL Ventures and an earlier seed round backed by Connecticut Innovations.
Official & Social
Watch: Demos & Interviews
Profile compiled from public sources including Wikipedia, Crunchbase, GlobeNewswire, Calcalist and company announcements. Financial figures are approximate.