Before Michael Malyuk had a company called HumanSignal, before its open-source software sat inside data teams around the world, and before venture investors put $30 million behind the business, there was a walk in the mountains. Malyuk and his future co-founders, Maxim Tkachenko and Nikolay Lyubimov, were trekking in the Himalayas, sometimes as high as 20,000 feet, and comparing notes from their previous engineering work. The scenery was grand. Their shared complaint was magnificently unglamorous: data-labeling software was bad.
Each of them had built a homegrown labeling tool. Each had seen machine-learning projects snag on poor training data and rigid interfaces. If a scientist needed to annotate an unusual combination of text, sound, image, or time series, the existing tool often dictated the work instead of serving it. The three engineers suspected the awkward little bottleneck was actually infrastructure. They came down from the mountains and spent long nights building what became Label Studio.
This is a useful place to meet Malyuk. His public biography is spare, but his trail of code, interviews, essays, and product decisions is unusually consistent. He likes the practical layer where an abstract ambition meets the inconvenient world. Artificial intelligence may be discussed in trillion-parameter poetry. Somebody still has to decide whether the thing in the photograph is a car, whether the sentence sounds angry, whether a chatbot answer is useful, and what to do when two qualified people disagree.
Build for the question mark
Heartex began in 2019 with Malyuk as CEO. The founding insight was not that labeling needed a prettier button. It was that data arrives with its own grammar. A tool designed only for single images would fail on audio. A workflow that suited a generic object might fail when the task required legal, scientific, or business expertise. Malyuk later said the answer to a data scientist wondering, “Does it support my use case?” always had to be yes.
That promise explains the shape of Label Studio: flexible interfaces, multiple data types, configurable workflows, and an open-source core. Open source was a product principle and a social contract. Users could adapt the software, inspect it, run it near their own data, and bring back edge cases the founders had never imagined. In a community AMA, Malyuk put the philosophy plainly: “As for open-source, it is democratization. We LOVE open source.” The capital letters are his. So is the enthusiasm.
The open-source bet created a loop. More users meant more strange workflows. More strange workflows demanded more flexibility. Better flexibility brought more users. By May 2022, Malyuk wrote that more than 100,000 people had used Label Studio. A community of thousands had gathered around it. The product was becoming a common surface where engineers, data scientists, annotators, and subject-matter experts could negotiate what their data meant.
A label is a decision wearing work clothes
Malyuk’s most revealing definition of labeling has little to do with boxes or tags. “The labeling process, in its essence, is a process of capturing decision making, creating a layer of knowledge on top of the raw dataset,” he wrote in 2022. Seen that way, an annotation is not clerical residue. It is a small, structured claim about the world.
The distinction matters because easy examples flatter a system. Drawing a box around an obvious car requires common knowledge. Judging the meaning of a contract clause, the quality of a customer-service reply, or the intent behind a sentence requires context. For Malyuk, the people closest to a domain should be able to encode what they know. He has argued that internal experts are the natural people to annotate and curate the training data in their field.
“The labeling process, in its essence, is a process of capturing decision making.”Michael Malyuk, 2022
He is equally candid about the messiness of those decisions. People make mistakes. Subjective tasks invite disagreement and bias. In his BentoML community session, Malyuk recommended combinations of ground-truth checks, multiple independent judgments, statistical error detection, and quality review. The point is not to pretend that one person supplies truth on demand. The point is to build a workflow that can notice disagreement, inspect it, and improve.
How a human signal becomes useful
This approach also gave Heartex a clean enterprise case. A flexible community product introduced practitioners to the tool. The paid platform added the control, collaboration, security, and scale larger organizations expected. In October 2020, Unusual Ventures led a $5 million seed investment. Seventeen months later, Redpoint led a $25 million Series A, with Unusual, Bow Capital, and Swift Ventures participating. Malyuk said the money would improve the product and grow the team.
Funding announcements can reduce founders to numbers, so it is worth noticing the less financial ambition he stated at the time: invest in the Label Studio community and put data-centric practices in more organizations. The business was selling software, but the operating thesis was about who gets to contribute knowledge. Domain experts were not an embarrassing human patch over an automated system. They were the people who could tell the system what mattered.
Heartex becomes HumanSignal
In June 2023, Heartex changed its name to HumanSignal. Corporate renames often arrive carrying a fog machine. This one clarified the argument. The company had begun with labeling, but Malyuk saw human feedback becoming relevant across the AI lifecycle: finding data, encoding knowledge, supervising models, evaluating results, and automating work without losing the criteria by which work is judged.
The rename also anticipated a shift in the technology. Large language models could generate fluent text, code, images, and conversation. Their outputs were less like a classification result and more like a performance. Evaluators needed to judge correctness, tone, relevance, safety, and domain fit. A thumbs-up button could begin the conversation, but it could not carry the whole burden.
Malyuk’s newer writing turns repeatedly to a deceptively simple question: what does “good” mean? He argues that evaluation matures in stages. A team begins with a directional reaction, writes a rubric, separates that rubric into dimensions, and revises it as products and expectations change. The complexity of the rubric, in his view, reflects how deeply the team understands its own problem.
That is a quietly demanding idea. It asks product managers, engineers, domain specialists, risk teams, and customers to convert instinct into language. They must decide where accuracy ends and taste begins. They must reveal conflicts that a single score would hide. A good rubric is part measurement system, part product strategy, and part peace treaty.
Data has to leave the desk
In November 2025, HumanSignal acquired Erud AI and launched a services division for multimodal data creation and expert annotation. The move expanded the company beyond software into operational work: collecting synchronized audio, video, motion, sensor, and other real-world data; recruiting specialist evaluators; and managing controlled environments for AI and robotics projects.
Malyuk summarized the premise in five words: “AI has outgrown the internet.” The web supplied extraordinary quantities of text and images, but systems that act in the physical world need experiences the web does not neatly contain. A robot must learn from motion and space. A new device may require synchronized sensor readings. A specialist model may need judgments from people with rare expertise. The next dataset may have to be designed, staged, captured, reviewed, and refined rather than found.
This is also a striking turn in the company’s business. In 2022, Malyuk distinguished Heartex from competitors by emphasizing software over outsourced labeling services. Three years later, the definition of useful data had changed enough to justify combining a platform with carefully managed operations. The through-line is not a fixed delivery model. It is the conviction that better AI depends on better systems for producing and evaluating information.
“AI has outgrown the internet.”Michael Malyuk, announcing HumanSignal Services in 2025
When agents act, reviewers need the whole scene
HumanSignal’s January 2026 evaluation engine announcement follows the same line into agentic software. An agent does not merely return a final answer. It chooses tools, creates intermediate steps, acts across several screens or systems, and leaves a trace. A reviewer looking only at the last sentence may miss the consequential mistake halfway through the journey.
Malyuk wants evaluation interfaces that fit these workflows instead of flattening them. Reviewers need relevant context, adaptable displays, and structured ways to record judgment across multiple kinds of output. The old flexibility question returns in a new form. Does the interface support this use case? Can the people who understand the work express what they see?
There is a founder’s lesson here, though Malyuk rarely packages it as one. Start with the nuisance everyone has learned to tolerate. Let practitioners reshape the tool. Build a community before the category has settled on its vocabulary. Then keep following the problem when its name changes.
Malyuk began with labels because labels were where knowledge met code. Today he talks about evaluation frameworks, agent traces, and real-world data factories. The vocabulary is larger, but the essential object remains surprisingly intimate: a person making a careful distinction, and software helping that distinction travel.