FOUNDER FOCUSIn this edition, Tech in Asia speaks with Ashish Dsa, co-founder and CTO of Arbor, a New York-based startup that runs anonymous AI-enabled voice interviews with frontline workers. Arbor's bet is that people will say more to a voice bot than to a survey or a superior. The interview has been edited for brevity and clarity.  What does Arbor actually do? Companies already run employee surveys, but they are usually one-to-five ratings, and people do not reveal much. There is fear around saying something negative. There is also hesitation around saying something positive, because people do not want to come across as condescending. We run anonymous AI voice interviews with employees, customers, and contractors. We begin with a consent screen, make clear that their response cannot be traced back to them, and then let them have a conversation. The AI can also sense when someone is uncomfortable. It knows when to move on instead of pushing them. That is not something a traditional survey can do. Voice AI is mostly associated with customer service and sales calls. What has changed to make it useful for employee research? Voice AI has only recently become good enough for this. Earlier systems worked like an assembly line: speech was converted to text, passed to a language model, and then turned back into speech. Each step added a pause. It did not feel like a conversation. Real-time models now combine those steps, so the AI can respond quickly enough to hold a natural exchange. That matters when you are asking someone about their workday. If the system pauses after every answer, people stop opening up. But the bigger shift is behavioral. People are getting more comfortable speaking to AI. They may still use it to book an appointment or solve a customer-service issue, but there is now an opportunity to use voice AI for something more open-ended: finding out what people really think. Why would people speak more openly to an AI than to a manager? People are more vulnerable with AI companions than with friends, sometimes even therapists. A human can judge you. With an AI, you feel you have a free pass. The setting matters too. In an unstructured conversation, people speak more. Someone who may have given a two-minute answer in a survey can speak for 20 or 30 minutes if they are guided politely. Trust also builds over time. In the first interview, someone may say a little. By the third or fourth, they are more likely to tell you what is actually going on. Can you give me a real-world example? We worked with a large freight operator in Los Angeles that had warehouses and terminals operating in isolation. It was hard for management to know what was going wrong across all of them. We interviewed workers across a few hundred terminals and found that the fastest teams had something very simple: free cans of WD-40, a lubricant oil. Their loading equipment would jam, and instead of waiting for a mechanic, workers would spray the hinges and keep going. That was one small detail from one conversation. But once the company made WD-40 available across its terminals, loading and unloading times came down. A consulting firm would have sent people terminal by terminal, and workers called into a room may not have said the same thing anyway. Does this risk turning employee feedback into workplace surveillance? That is everything for us. If people think we can trace a response back to them, they will stop talking. We do not build a system for continuous listening. That would be an anonymity nightmare. We use selective listening: a worker is invited to take an interview and chooses whether to participate. There are exceptions for serious safety issues. If someone indicates that they may harm another person, the company gets an alert. But even then, it does not get a name. What does the technology look like behind the scenes? For these conversations to feel natural, latency matters a lot. If the AI pauses too long after every response, it stops feeling like a conversation and becomes a monologue. We often use real-time models, where speech comes in and a voice response comes out through one system. That gives us less flexibility than stitching together different speech and AI models, but it is much faster. Where does Arbor go from here? Most AI is still one-shot. You ask a question, it answers, and that is it. We want to build something that gets better with every interview, correction, and piece of feedback. The people at the top of a company often do not know what is happening on the ground. The people on the ground usually do. The question is whether companies can finally hear them. |