FOUNDER FOCUSWe speak with AI founders and executives about AI, startups, and whatever else is on their minds. In this edition, Tech in Asia sits down with Dahua Lin, co-founder and chief scientist of SenseTime, as the company recently launched SenseNova U1-Pro, its flagship image generation model. The Hong Kong-listed firm was once the world's most valuable AI startup, built on facial recognition technology that made it a global leader in computer vision. That same technology put it under US sanctions over alleged surveillance of Uyghurs in Xinjiang, which the company denies, and the generative AI wave left it playing catch-up. Now, SenseTime is betting on multimodal models: systems that handle language and vision together in a single model rather than treating them as separate problems. This interview has been edited for brevity and clarity.  SenseTime open-sourced its U1 Lite image model in April. What can U1-Pro do that the lightweight version couldn't? The lightweight version generates pretty images, but its ability to control complex details is not as good. U1-Pro can produce professional-grade posters and infographics with very complex information inside, at a level where you can go directly to the printing shop, print it, and hand it to your customer. We have 200 designers in our team, and every image produced by the model is assessed by several designers to determine whether it is ready for production. Over 70% of the images produced by U1-Pro are production-ready in the eyes of a designer. For the lightweight version, it is less than 20%. U1 Lite is good for consumer or family use, while U1-Pro is aimed at commercial use. Why did you decide to keep U1-Pro closed source? We will continue along both directions. We will keep open-sourcing our lightweight models, and we will open-source U1.5 in the not-too-distant future. U1-Pro is our flagship, and we will monetize it through an API with subscription or usage-based fees. With U1 Lite, we wanted to promote awareness of unified multimodal models and give the community a basis to explore. U1-Pro is a much bigger model, and its deployment involves an agentic generation loop. Producing each poster takes multiple steps: a design plan, decomposition into steps, generation of different parts, then several rounds of refinement. There is some secret in how we compose that trajectory to meet aesthetic criteria, and that is important for commercialization. GPT-Image 2 and Nano Banana are also proprietary. When customers use our API, we can directly see how they use the model and iterate quickly. That is more efficient than open-source feedback. SenseTime was once the world's most valuable AI startup. Is this model, and the multimodal approach, the path back for your share price? A stock price in the AI landscape can go up 10x in two months and be cut to a tenth within weeks. These ups and downs come and go. That is not what I consider the most important part. What matters is our vision, which is to achieve AGI in the physical world. Language models alone are not enough for that. An agent must not only reason in the digital world but also interact with the physical world. It needs to see the world, then think about it. So we unify images and language in a single model, so that logical reasoning, spatial structure, and understanding of the visual world are combined into a single brain. That is what you need for a robot that works smartly in the physical world, and once that happens, it opens up a trillion-dollar market, not just for SenseTime but for every company. That is a long way off. So we open up commercial sectors one by one along the way. With U1-Pro, we are opening up content creation and design, and that provides financial support to continue the path. This model does a lot of reasoning, which means a lot of compute. How are you handling US export controls on chips? We have a very diverse portfolio of suppliers for our computing power. This model can run on any kind of GPU, including those from Nvidia and those from vendors in China. The unified architecture is very concise, which makes it easy to adapt to different chips. We have tested it internally on different kinds of chips and it runs well on all of them. We don't have a big problem supplying the power for inference. How important is U1-Pro to SenseTime's top and bottom line? This is a very serious attempt to open up another commercial sector. AI coding has already been proven a profitable business, and people are willing to pay for it. But coding is not the only sector that can benefit from AI. We deeply believe content creation and design is the next one. It is difficult to anticipate numbers because we are yet to launch. But if a customer is willing to pay, say, US$100 to employ a designer for a poster, it should be a no-brainer to pay one tenth of that for an AI tool that produces the same quality. In terms of image generation, what gap do you think OpenAI and Google leave in Southeast Asia that SenseTime can fill? SenseTime has been the number one computer vision solution provider for 10 consecutive years, not just in China but globally. Many customers need not only LLMs but vision, and a way to combine both with their specific application scenarios. Google or OpenAI can provide components, but they are not as interested as SenseTime in serving those customers closely. We provide the entire solution, and we grow with the customer. We already have key accounts testing U1-Pro in mainland China and Hong Kong. I am based in Hong Kong and travel to Singapore and Malaysia a few times a year. From those markets, I see a lot of demand for multimodal AI. Asia remains the market of highest priority in our commercial strategy. Compute costs are squeezing gross margins across the industry. How do you see that playing out? In the long run, it is not about token economics. Selling tokens will be under high pressure as computing costs go up. Our view is that eventually, you want to sell the results. You won't be counting the tokens consumed in the process - you'll just pay for the poster. Right now, each U1-Pro image costs several US dollars to produce and takes around five to 10 minutes because it is a long-horizon task with multiple steps, while U1 Lite sometimes takes just seconds. The exact pricing is yet to be determined. If we offer a competitive price while producing at a much lower cost, we have a large margin. We have our AI infra department, and we built Asia's largest computing center, so we have a very strong team optimizing our compute. Together with optimization of our model architecture, this will substantially bring down the cost. |