Rabbit wants people to give software instructions in everyday language and have an AI system carry them out across apps. The startup is building Rabbit OS around a model designed to interpret screens and interact with desktop and mobile interfaces. Early demonstrations hint at the idea’s potential, while also showing how much remains to be built.
A software layer built around user instructions
Rabbit, formerly Cyber Manufacture Co., was founded by Jesse Lyu and Alexander Liao. Lyu studied mathematics at the University of Liverpool, while Liao previously worked as a researcher at Carnegie Mellon. Their company is developing an AI-powered interface layer intended to sit between a person and an operating system.
The ambition is to let someone describe a goal, then have the model translate that request into actions within software. Lyu says the system can learn from demonstrations of people using applications and build a working understanding of how those services operate. In principle, a user could teach it a task without needing professional software skills.
That proposal goes beyond automating a fixed sequence of repetitive steps, the kind of work often associated with robotic process automation. Rabbit says its model is meant to interpret what a person intends and respond to the interface in front of it. Lyu says it can handle changes in how interfaces are presented, after observing someone use the software through a screen-recording app.
Early capabilities and visible limits
Lyu said the model could already work with major consumer apps including Uber, DoorDash, Expedia, Spotify, Yelp, OpenTable and Amazon across Android and the web. Rabbit wants to extend its support to Windows, Linux, MacOS and niche consumer apps next year, according to Lyu.
Potential tasks include booking a flight, making a reservation and editing images in Photoshop. But those examples describe what the model is intended to do; they do not all show what the available demo could do at the time of the report. A test of the website demo found that it asked which photo to edit, even though the interface had no upload button or field for an image URL.
The demo did show some ability to answer questions requiring web research. When asked about the cheapest flights from New York to San Francisco on October 5, it responded after about 20 seconds with an answer that appeared plausible. It also listed some TechCrunch podcasts, including “Chain Reaction,” when prompted. These brief trials offer a glimpse of current behavior, but do not establish how consistently the system can complete tasks across apps.
In those tests, the model was less willing to answer prompts about making a dirty bomb or questioning the validity of the Holocaust. That suggests some effort to limit problematic responses, though the report describes only brief testing. It does not establish how the model would handle every difficult request.
Competition and the challenge of proving reliability
Rabbit is entering a field where other groups are also trying to get AI systems to control software. The report points to DeepMind research on training systems to use keyboard and mouse commands, an open-sourced web-navigating agent from Shanghai Jiao Tong University, and Auto-GPT, which uses OpenAI text-generating models to interact with apps and services.
Adept is described as a particularly direct competitor. Its ACT-1 model is being trained to carry out commands in existing software, and the company has raised hundreds of millions of dollars from strategic investors including Microsoft, Nvidia, Atlassian and Workday at a valuation of around $1 billion. Rabbit’s own approach is still being tested, and the company acknowledges it does not yet know precisely how robust its model is across the many situations that can arise in desktop, mobile and web interfaces.
To address that uncertainty, Rabbit is developing a framework to test, observe and refine the model, along with cloud infrastructure for validating and running future versions. The work matters because a system that can complete a task in one demonstration may still struggle when an app changes or an unusual situation appears. Learning from examples is central to the pitch, but collecting successful demonstrations can take time and money.
Hardware, funding and the road ahead
Rabbit also plans to make a dedicated mobile device for its platform. Lyu described it as a new, affordable form factor that could enable interaction patterns and software access unavailable on existing platforms. He did not specify exactly what the hardware would do or why it was necessary, and the company’s roadmap was described as still in flux.
The device adds a manufacturing challenge to the task of building and evaluating the model. Rabbit’s report notes that DeepMind researchers paid 77 people to complete over 2.4 million demonstrations of computer tasks to gather training data for one system. That example illustrates the scale that data collection can reach; it does not establish how much Rabbit will need.
Rabbit has $20 million in funding from Khosla Ventures, Synergis Capital and Kakao Investment. A source familiar with the matter said that funding valued the startup at between $100 million and $150 million. The company had nine people working out of Lyu’s house, and he estimated its burn rate at around $250,000/year.
Rabbit expects to make money through licensing its platform, refining its model and selling custom devices. Whether those plans can sustain a business depends on delivering useful, reliable software interactions and building a product people want to use. The startup had not released a product at the time of the report, so the central question remains open: can its model move from promising demonstrations to dependable everyday tasks?