Transcript: Kb41Dtlx1Uc
Source Video
Local Cache
raw/sources/youtube-transcripts/KB41dTlX1Uc.txt- 9,219 words
Transcript
[music] Can you guys hear us? >> Sound check. All right, give it up for local AI, everyone. [applause] >> I hope you guys are excited as we are. Um, this is woo. This is the local AI summit. So, we're going to be here all day talking about local AI. And the reason why is we hit an inflection point this year. Not only did the models get really good, but the harnesses got really good. And this happened really fast. It's been, I think, a struggle for anyone here to keep up. I felt that most when I saw one of Andre Karpathy's tweets in November. He tweeted that you can't really trust these coding agents alone yet. You have to monitor them with an eye like a hawk. Three months later, he tweets that he's struggling to keep up with the capabilities of how good this all has gotten. And the thing is, both times he was right. This space is progressing really quickly. And it's wild how much you can do. And honestly, the only thing to do is just try to use it a little bit more today than you did yesterday. That's how to keep up with this space. And that's exactly what you guys are doing right here. And so I'm really excited about this panel. The way that we use AI has also changed. I'm not just using chat bots. I'm not just asking simple questions. When we got reasoning models, the the profile of how the AI model responded changed. Not only is it bursty and responding to me, but before the burst, it kind of plateaus for a bit. It's reasoning. It's turnurning on tokens that I'm not consuming. And then we got agents. And suddenly I don't even want these agents to turn off. I want these always on agents. They can always be productive if we set them up. And so we have enterprises that want to put a lot of their IP into this because it becomes more useful. We have consumers with the same thing. I want to give it my health data, my medical records. I want to give it footage from my home camera. And both enterprises and consumers, we don't want that stuff to leak. And as you also have a profile of uh tokens continuously generating, suddenly costs matter. And so local is amazing for both of those things. you get to make sure that you are plateaued on the costs uh for the tokens that you're generating and also uh everything sits in that room. So we have amazing demos here after these talks where the everything that's being run stays on those devices. It stays in this room and so that's a really nice guarantee. So as we turn over to the panels um first do you guys want to introduce yourselves? >> Yeah. Should I start please? >> You got the mic? >> Yeah, you're miked up. >> So yeah, I'm Alex. I'm the co-founder and CEO of ExoLabs and also the creator of local.ai. So we're we've been working on local AI for over two years which feels like a lot longer in this space and you know from the days of you know running Llama for AB on two MacBooks to now where we have you know demo there of Neimatron ultra running on four sparks you know our mission is to make AI more accessible and you know it's crazy to see what is possible now um >> cool hey everybody I'm Matt um I make content and uh videos. We have a newsletter all about artificial intelligence. I'm an AI enthusiast. Uh and yeah, thanks for having me. >> Hey everyone. Um my name is Ahmed Osman. Uh I am the founder and CEO of Osmantic. Uh somebody jumped into my DMs asked, "Hey, what does OS man stand for? Open source man." So that became a company. [laughter] And uh I also moderate local lamb subreddit. I have been in this local AI space since 2022. And uh yeah, let's uh let's make open source and local AI one. >> Love it. >> Open source man gets the round of applause. [laughter] >> That's appropriate. Today is uh I think a pretty momentous day for the panel. Um you know, Fable just came back, but it's a good reminder of why we need to have access to Frontier Intelligence to be able to build anything. Uh I'm Joseph. I'm co-founder CEO at Roboflow. We do all things vision. I kind of like to joke that vision is like the original local AI because like everything needs to run probably concurrent with like low latency alongside your video. Uh where the images and data is being captured. Uh looking forward to the discussion today and all things why local AI is the future of making AI useful for everyone. Everyone let's give Joseph and everyone a round of applause. So one thing that I'm really excited about for all the panelists is you guys were all early to the space in your own way. Uh and you know very early and I think as we all feel this inflection point now uh when do you guys feel like you felt the inflection point? >> Yeah I can go first. Sure. I mean look when I first saw Llama right I got really excited the idea that I can actually download this intelligence run it from my local computer. It was very exciting. I'm a tinkerer at heart. I'm a builder at heart. I I've overclocked PCs for years. So, this just felt like so good. So, felt so interesting to be able to have this crazy alien intelligence running on on my device in my office. Uh and that that was the first time where I didn't even think it was possible to be on till I saw Llama and I was like, "Oh, wow. This is incredible." Um so, that was really the point at which I was very turned on to it. >> I I think I um it's kind of the same for me. Uh yeah, it's Lamatu. I think that um finally on my 4090 RTX 4090, I'm like, "Wow, I can actually understand this black box and customize it and play with parameters and like you know, sampling parameters and all those uh configurations and see how inference engines work." Uh that that that made me feel like you know something clicked. It's like oh magic. I can be a wizard now with this thing. I can control how it how it plays to my ways of thinking. Um and from there it was you know the rest is history. Basically, I've been very vocal about why local and open source AI must coexist with cloud and how how we are supposed to like, you know, push for that and try to understand it as much as possible and teach people about it. >> Totally. I actually felt this as well. I was on a plane. I think it was in 2023 and I had uh a model running on my phone and it was awful. It took 20 minutes to complete a sentence, but I was I didn't have internet and so I was able to like ask uh essentially intelligence something there in the plane and that was my first feeling of like hm this is going to be something really cool. >> Do you know that you can run the equivalent of GT40 on your iPhone now? >> Really? >> Yeah. It's quen 3.5 the four parameter for billion parameters. It it's basically the same quality as something that used to be served in data centers and it's in a device in your pocket. Yeah, that's crazy. >> Yeah, I think if you just think about like how crazy that GPT40 moment was and now you can run it on the phone, it's like I think a lot of this stuff was just, you know, you had to see the vision of like where things were going because it was a toy, right? It was like, you know, the the first I think there were a few key points for me. One was like running Llama 45b locally, right? That was like the first really big open source model. And I remember like the gap it closed the gap quite significantly with the frontier closed models and open uh but it ran at two tokens a second so it wasn't useful right um and then you know another big moment was deepseek um both v3 and r1 deepseek v3 was like a massive so like you know llama 4 was dense so it's super slow seemed to unlock the performance so it was like oh wow you know with the devices I already have like with a you know a Mac uh Mac studio or Spark, I can run, you know, this massive model at actually like decent performance which is comparable with, you know, what you can run in the cloud. And then uh I think just recently to me GLM 5.2 is a big moment as well. Uh because again it's closing the gap and it's like you know Opus level um and you know we have it running over there on a device that like literally can fit on your desk uh the DJX station. So this is to me just like you know a trend and it's going to keep increasing right like there's going to be smaller and smaller devices with less memory better compression we'll be able to run more capable models locally and you know soon it'll be the default >> I'll keep with the uh the theme of of airplane stories. So yeah, I I agree that there's been like multiple moments over over time of what's been what's been going on. But one person >> I have a presentation for that after this. So stay around, please. >> All right. Uh one example moment where this like came up for me was uh I was on a plane. I was sitting next to someone man this is maybe two three years ago and they were hard of sight and they were using if you on Apple you know if you take a photo of something you can use the accessibility settings and it'll describe the photo to you >> and we were seated on the plane together and they were taking constant photos to understand where we seated and how to get buckled in and these sorts of things and I remember they took a photo of the seat back in front of them and the Apple accessibility described the photo described it as like a printer or something and it was like oh yeah like you're seated like by a printer and of course the individual knew from the context that couldn't be right. Uh, and I was like, you know, I wonder Lava had just come out, a multimodal model that like kind of just described things with like the pivot to text base rather than like natively multimodal, but still kind of a useful example. And I was curious. I was like, well, let's see how Lava would do on the same thing. So, I took a photo of like the seat tray in front of me. And it aptly described it. It's like, hey, you're on an airplane. That's the seat tray. And I showed the person sitting next to me who had never seen local models and never experienced local AI. And what that stood out to me is I was like, man, a company that's a trillion dollar plus market cap business, shipping the latest intelligence on their phones for describing visual settings is inferior to something that's broadly accessible and available to anyone. And that was such like a watershed clear moment that like even the largest companies don't have a monopoly on the frontier of intelligence. And so increasingly making that be accessible to others um is I think going to be critical for understanding all the massive impacts the tech will have. >> And that was lava. Yeah. [laughter] >> Yeah. It's interesting to hear >> that's like 2022 2023, right? >> Around there. Yeah. Yeah. That was a while ago. It's interesting to hear you say about the the accessibility of that, right? The uh taking this Frontier intelligence, what made it useful was giving it peripherals, in this case, a camera so that it could access the data in front of you. And I think that's where the inflection point this year was so much more than just models, but also these harnesses and what you can give it access to. It was CLI so that you can give business systems uh and plug those right into your agent. It was uh you know 40 was really exciting because I was coding with it, right? I was taking code snippets from my codebase and then I'd go to chatgpt and I'd paste it. I'd give it as much context as I thought I should and then I would take the result and I'd go back to my codebase and paste it. And what was really great about tools like cursor was it essentially was this harness that said what if I just had the full file system and I can let these agents reason on what files they needed and so on. And so the way that you know you get more use out of all of these local uh out of all of these local models and systems is how does it how can it interact with the real world. I'm curious um Joseph so you guys got the start in vision and it seems like vision had to learn the hard way what a lot of LLM are dealing uh discovering right now. What do you think is like a big lesson that the language world is discovering right now that uh we could look to vision for? I think um one of the things that language has an advantage of is it's inherently a human construct. So language generally exists where people exist and that also means usually you can use uh infinite amounts of of computer what is available in in a data center whereas a lot of vision uh you're computed. You're running maybe where there's low internet connectivity. You're running on a device. You're running on a robot and the amount of compute you have available to you is is what um is what you get. And so what did that do? So that created I think an emphasis on specialized learners faster >> because it's like okay I need this model to run in a limited context um and I'm not going to prioritize full world scale generalizability instead I'm going to have specialized in domain context work and what's interesting is I think you're actually seeing language follow a similar thing you're describing in the context of like harnesses for a given context um but you know tools like uh any sort of coding agent where you have a really good harness and it's specialized for doing those sorts of tasks but you're also seeing that even in like you know tax preparation or legal preparation of last mile fine-tuning adaptation specialized learners and so in some ways I feel like the pendulum swinging back to actually having specialized models even in language just as much as vision I think is one thing um and then you know there's a whole different thing around how you get the most out of compute optimization running on on a niche device >> um but I know that probably the exo folks can talk speak more >> absolutely >> um but yeah I think specialized models I mean, you remember there was like one model will rule them all world was kind of thought and I feel like the pendulum is swung back to people realizing special >> specialized models. I think we definitely at NVIDIA we see that it's going to be a multimodel world. We definitely agree with that. Um there's actually one of the panels that's coming on later I think at 3:20 p.m. is all about model routing and so that's that's really exciting. I guess as an enthusiast uh and as you use some of this is that uh how your usage pattern looks? Are you using many models? What does that kind of look like? >> Yeah, absolutely. everything from obviously like the top models Fable uh to kind of the more workhorse models with Sonnet local models for things that maybe don't depend as much on low latency. Um yeah, like the multimodel world seems like a no-brainer at this point. uh especially as you're seeing all these enterprise companies come out and look at their budgets and and think, okay, well, I I need to continue to increase the the total number of tokens that I'm consuming as a company, but I also don't want to just completely blow out my budget. I know Coinbase and Brian Armstrong just came out with that just great post the other day talking about uh how their tokens are are exploding yet their costs are staying flat. And and that is because they're using a mixture of different models. You don't need the top model for every single use case and in fact most use cases you don't uh I think the most obvious application is let the top model plan uh the the architecture whatever the kind of top level plan is and then the actual execution of the code can go to uh a more reasonably priced smaller model. >> Your most intelligent should provide you with the overall plan and then subtasks for your smaller executioner like executioner models and that's exactly the future. >> Yeah. And especially like you know we're talking about local um local models are great at writing code but maybe we offload the actual top level planning to one of the frontier models and you you save a bunch you control more of the workflow. It's it's a it's a nice pattern. >> Yeah. I I think the market wants this. So it wasn't clear like a few years ago when this stuff was just starting to play out, you know, would there just be one big model that everyone is using? But clearly that is not what people want. That is not what enterprises want. They don't want to be told what they can do by Dario. They don't want to be paying for the same model for all their workloads when some some workloads don't actually need, you know, a giant gigantic model that costs $50 per million tokens. Um, they want control, they want sovereignty, they want the ability to switch out models, they don't want to get rugpulled, you know, uh, from one day to the other, uh, because of some safety risk or whatever. Um, and you know, that's really what's driving a lot of this progress. So, you know, I'm more optimistic than ever, I think, about sort of, you know, where things are going with local. I think that, you know, the market is basically pulling a lot of this stuff um out of, you know, startups, out of enterprises are building solutions around this and, you know, it's it's yeah, it's it's happening, right? >> Totally. I feel like this is a new frontier where as we go multimodel, that becomes really difficult. Of course, you need to route between the models, but then you have to figure out how to provide necessary context to whichever model you're routing to. If you were making a plan and then breaking it off into sub aents, what is the framework to do? So, how should you do this? These are all open problems. And, you know, I think just being playful is uh kind of the best way to to approach this. But I like the way that you worded that of like uh the market's kind of pulling for these answers. It's reaching in this direction. We're looking for startups and companies to fill this gap. Um, makes a ton of sense. You know what's funny? Um, as any any enterprise consumes any any piece of software, it matters a lot. If you're going to build a foundation, you want to know what the versions are. You want to know what your model is, which whether that version changed in order to see change behavior later. So having uh also more control over all of this so that you know exactly what version it is, it can't be changed. Uh not even because of any sort of you know regulation, but just simply if there are updates, you choose when you're opting into whatever it is on your entire stack. >> Yeah. Uh so [clears throat] when you think about um this wave that we're seeing here of like this civilizational infrastructure that is called AI you have to consider the the potential of you know things being taken away from you and your sovereignty and uh how can you be in charge of you know this thing full stack end to end that's hardware software and everything in between That's the model weights. That's the specific version of the model as you were saying. That's how you fine-tune it. And picking up on Jos's you know thing um statement here [clears throat] about small and specialized models. I have been like, you know, I have been a proponent of that for quite some time. I I have tweets about that in 2024 saying small and specialized models are the future and it is going to be very user by per use case per workflow for businesses you know to to decide on their niche domains and you need to start on boarding them now because you want to collect the data points that races that's what you're asking about N you're asking about how do we get there how do we decide which model gets routed to how do we decide which use case gets goes to which model or which endpoint you You need to collect data and you need to be ready to basically break down um the traces that you've collected from your employees from everybody in your organization and decide how can we make the most optimal use case of these data points to which models and collect feedback as well and you can actually automate that with agents as well. So like this is like something that it's you know the the whole thing about RSI right now recursive self-improvement it also applies to even agents and harnesses and to to workflows and to use cases and to enterprises. You want to describe that for everyone too maybe for folks who don't know RSI >> uh recursive self-improvement it means that uh a model would basically rent its own compute and then start training its own checkpoints and then you know uh deploy it its next version so that it gets updated in certain ways change its behavior in certain ways so basically a model is training itself >> totally for our product brev which just makes it really easy to get a GPU we've been focusing more on agents as a first class audience because we're seeing this so we're seeing growing usage from agents that just want to go grab a GPU uh directly. >> Exactly. >> So yeah, you know, it's funny. So the second thing you talked about was uh not just multimodel but then all these optimizations to actually run it performant when you have less compute than in a say data center. I think Alex this is some fun work that we did. So um Alex has actually set up a second headquarters inside of Nvidia. we got a conference room and uh we were we were having dinner and he mentioned that he really was motivated to get to squeeze out every drop of performance we could on the DGX Spark and so we said hey let's get a conference room at Nvidia come bring your team and anytime you need an expert across any pillar let's just go and pull that person into the room and yeah >> I I actually have a follow-up here uh so since I have been part of local lama for quite some time um when I think about home labers I think about the individuals that are actually are within enterprises on boardrooms making decisions about you know the models that we're going to host or you know where we're going to get our AI from. Uh the things that these home labers focus on are optimizations. It's basically how can I extract the most economical value out of the hardware the constraints that I have the software that I have. You know quantizations came to be because of that. When we're thinking about enterprises it's the same thing. It's constraints. It's budget limitations. It's how can I make the most use of the number of tokens that I can get under the hardware that I have. Whether that's a DGX station, you know, um centrally in a small medalsized business or, you know, a DGXP300 cluster, it's all about how can I optimize the software for the latest and greatest frontier opensource model out there and get as much value as possible out of it under the economical constraints that we have. I think it's it's crucial to see home labers and that's why you know this is local AI it's local AI but it could it could be on premises it could be colloccated hardware it could be rented clusters it could be from brev basically it's just the idea about controlling this thing end to end from hardware to software to model endpoints to model weights and you know collecting all of that data to train your specialized and smaller models that will be more efficient as as you go. Um yeah, the future is great and the future is local. [laughter] >> Yeah, absolutely. Um you know, making sure it's your dates, your your weights, your compute. I think one thing that was really interesting, we did end up getting 10x performant improvements on the DGX Spark and we sent an update to Jensen, to the executive staff, to the teams that were helping. And one line I really liked in that email was that we didn't solve any new computer science to do this. We actually took things that the experts at NVIDIA had already solved and was out there. And I think what we worked together to do really nicely was assemble it in a bouquet. And I think that speaks to some of the the usability here, right? We're talking about what is capable. What what are the capabilities? But do you feel like it's just capabilities that are holding people back from adopting this? >> Yeah. So just to tell the story a little bit what happened here. So I think we had dinner on Thursday of the week and then you know an email got sent out on Friday. So this idea came up like why don't we do a lab, right? Why don't we like you know this to me sounded crazy at the time. I didn't think this would be possible but it was like you know let like Nata was like oh let's like get a bunch of people from Nvidia to help you guys and work with you guys to improve the performance on the spark right and I was like okay like let yeah let's I mean if we can that would be great. Email got sent out I think pretty much that night and then on Friday Nata told me be here on Monday um at NVIDIA HQ. uh we're going to have you know teams of people here that are going to work together with you on this and you know turned up on Monday and uh Jensen talks about this concept of swarming. Uh it's you know basically the idea that uh the the whole company will like mobilize around something. >> Have you ever seen little kids play soccer? There's a ball and everyone just attacks it. It's like a tackle. It's just the mit of the soccer ball. >> Yeah. And I had heard about this idea, but like I'd never experienced it. And I can tell you like we were just back toback that whole day. People coming in and out of the room, you know, from all various teams like the Neatron team, um, you know, people working in in on data center stuff, people working um on VLM. I I realized, you know, Nvidia has a team for everything. So there's like a VLM for Spark team which is oddly specific but like you know they have all these teams and you know we basically got to pull those resources in and you know you know in those three weeks or so as you mentioned like we basically got 10x performance versus you know what um Nvidia had running on the spark um in their existing playbook which was using Hermes agent. So we did a bunch of optimizations there um using VLM um as the the sort of inference back end there. Uh doing a lot of work with the like tuning the models like quantizing the models uh to you know be fit for local. So what you'll find is I think one of the great things Nvidia has done with the local hardware is it's the same architecture that's running in the data center as running on the spark. So it's Grace Blackwell. Um so the hardware is like fundamentally the same meaning you actually get a lot of things for free. So for example like the kernels are already really good. However there is like a lot of tuning and a lot of like configuration that is right now designed specifically for the data center. So a lot of the work that we did was not like inventing anything new but it was actually just tweaking things to work uh more performantly on the spark. And you know the hardware is extremely capable. It's like you know it is data center level hardware that's literally can sit on your desk. Um so it's just about like you know how do we activate that and I think this has been one of my learnings from the last few years is like we already have the hardware like the hardware is already really good and you know the models are getting better. Um they're getting a lot better at compression so you can fit more on a smaller device. So like you know I think it really is just a matter of more people looking at the space and working in it and you know we will be able to do amazing things like we have you know neatron 3 ultra it's a 550 billion parameter model running on four sparks over there 30 tokens per second so like >> that's that's huge I want to say something here like it's it's so great to see the community being proactive I think of Alex Gmail right here who's exo as somebody who's member of the community who's trying to push the frontier opensource intelligence to the next level. I love how Nvidia is working with them on that. When I think about, you know, this space, I think we're in the '9s of the Linux operating system and we are like just starting. The infrastructure is not there yet. We need so much more. Uh I was even telling Alex, I want to see how we can collaborate. We have like an open source uh uh deployment system called ODS that is a bunch of open source tools that we deploy in each hardware. It configures stuff, configures agents and basically uh gets you set like you know get gets you set up with the entire infrastructure end to end that you need for your agents locally. This kind of stuff we need that optimized for every piece of hardware out there so that we can get more and more people on boarded have the open source adoption you know just go to the moon. That's what I want right now. This is why we're here. We want everybody to know that local local and open source AI can run on anything starting from your phone to your DJX Spark as you're saying Alex to DJX stations to the next level in the data centers and it will deliver you frontier level intelligence. >> Yeah, absolutely. I think this is a really good question for you Matt. You know you test everything for a mainstream audience. You're hearing about all of this technology. You're using this aggressively. Where does it still fall short for the uh average user for AI? So I I think about [clears throat] two things. The average kind of personal user then the average enterprise using uh open source. I mean it's it it needs to basically be as simple as opening cursor. >> It needs to be maybe slightly more complicated than that or slightly more complicated than just installing codecs and opening it up. Right now it to be fair it is quite far from that. it like the the stuff these guys are doing is incredible, but it is more sophisticated than what most people including myself are going to be capable of. Um, you know, let alone a business. Uh, I think >> a full-time job to just do this kind of stuff. That's why we need to automate it. >> And it it really does need to be point-and-click. And once it gets there, and there's there's a lot of great open source projects, there's a lot of great projects in general that are getting there, but we're still not quite there. Uh the other thing to make it widely adopted is to allow people to better understand what use cases are appropriate for what what type of model for what type of harness what type of hardware knowing exactly the use case that I can get out of my home system or something that I'm I'm you know renting from uh service center uh that that is incredibly important as well >> and I I think it shouldn't be just in documentation. I think that's where open source becomes so difficult for people. I think it needs to be a point and click and it figures it out on its own. That's what ODS is about. That's what I think EXO is about. That's what you know as we grow more and more and building this infrastructure, we need to be thinking about the user experience for your average everyday user, not us as technical folks here because this is AI engineering um you know summit and I'm pretty sure that we all can manage our way around this stuff. But your chat GBT users, your cloud users, whatever out there, they want this to be an alternative that is just click play, send a message, use an agent, and done. >> Yeah. Most most people really don't want to know about the details. They they just want it to work. And even if it worked seamlessly within cursor or codeex or any of these other things and it just worked and they didn't have to think about it, that is number one uh that for for the vast majority of people. So [clears throat] interesting on that in a multimodel world where we have to pick a model do you think that that's something that is a product um as we're or is that is that plumbing is that something >> matter you talked about the harness it's exactly that it should be some like whatever the kind of the front entry point to somebody's AI experience is that should be what is choosing which model at the right time for the right use case this is a very difficult problem there are you know open source projects closed source uh companies building routing That's only one piece of it, right? So, go ahead. >> No, like Carter from Nvidia, he's he's going to be a moderator here. Hey, Carter. Uh, looking good, man. So, um, he I I I shared ODS with him and he told me, I love how like it immediately downloaded that two billion parameter models allowed me to start playing with it and then it started downloading the next model that would work perfectly on on my device. That's the kind of experience we need to be giving these users when they we're on boarding them. Don't make them sit down and have to think about all these quantizations and all these extensions and all these weird things. That is too much work for your average user. That's how we lose them. >> Mhm. Yeah. It's interesting, right? Like the when you're not having to deal with multimodel and you have just this generally good but big model, you can kind of ask anything. You don't have to be as sophisticated or you don't have to understand your use case too well. But as we talk about specialized models, if you're going to specialize in something, you need to know what that something is. And so the amount that I mean can you guys speak to maybe what it's like to move into a specialized model or to build one or have you have you >> I can start with that. Uh on the cloud uh you're basically getting the normal distribution. So every time that they are training something they are taking the average feedback from everybody when they are being happy or sad or like you know um they are okay with the answer or hey no this is not what I meant go back or change the model and answer with a different model. That's the kind of feedback that allows you, you know, cloud providers to fine-tune the next model. But if we're talking about small and specialized models for use cases, that's a lot of compute to train models. That's a lot of um storage for these weights. Uh that does not happen without each use case or each business entity etc. focusing on their own um patterns and use cases and like how they handle agents and how their employees uh message these agents and what workflows they're interested in. So it's really you know it's per use case it's per entity it it's not something that we can just generalize and that's why you know continual learning is not being talked about by the frontier model uh as much it's coming though it will happen and it needs to be running on local hardware for it to happen. That's how you can have something that is so optimized that is not just plotted markdown files sitting in one rebel and you know you think that that agent is not going to lose track of which skill to update or what to edit or which memories to update because at some point context lens becomes inefficient and that's one issue that's the current paradigm of agents is basically just saving to markdowns the next one will be updating the weights and that needs to happen locally. Yeah, there's it's funny too because you know uh there was a tweet I think by Swix uh that was like why has fine-tuning as a service not taken off. It was like a couple months ago. Um and I think a big reason for that is that model customization itself is also a very hard problem. And then what models can you hack on? Uh that's I mean that's why Nvidia releases Neotron. It's a open source model but everything from the dates sorry the data to the weights to the uh the recipes for how to do so and then of course you know the final model everything is open sourced um just so that it is a model that you know you can safely uh use and customize >> you know it's interesting I keep kind of flip-flopping on this point I think people look at how good the generalized models are and you give it the right context is that going to be better than than having a fine-tuned model and all the work that comes with that And I I keep going back and forth because there is a lot of value in being able to have kind of a smaller very specialized model and and maybe a bunch of them working in unison to accomplish whatever task you have. Um but I I yeah I'm not sure what what do you what do you guys think about that? Um, if I can chime in on this one, like I think the the beauty of like the open source ecosystem is that like all of these paths kind of get to be explored and then whatever wins like you know it's it it's like what happened with speculative decoding um you know there's all these like different ways of doing speculative decoding which is this idea that you can use a smaller model to basically approximate a larger model to speed it up. Um and you know you've seen like recently literally I think in the last week there's been like three different like quite you know seem to be breakthroughs in like speculative decoding that have just come from like different places. One from deepseek um there was some work that was done by like modal and uh the SG Lang team to like build uh Dlash draft models for like various Quen models that are like a big improvement on the previous ones. um there's like all this work being done and so like we kind kind of get to see like all of it and you know basically like whatever ends up being the best thing will just be the thing that kind of right. >> That's a great point where maybe it's not users and consumers actually customizing their own models and using the specialized models but someone who has a need does so and does so in an open source fashion so that someone else can just adopt it. Right? If I'm going to start to do image gen for a particular use case I can just go find a model. Actually we saw this a lot also with like early Lauras on like a lot of image gen models. I feel like that was really big as you know in the 40 era as well. I think ultimately what the end user cares about is does it solve my use case and can it do so within my budget >> and if there's another service that can manage the entire fine-tuning process and and I don't have to think I know I keep saying I don't have to think about it sometimes I do think about things but like in this specific use case um >> thinking about too many things already to add one more thing to >> like I I care about my business I don't want to think about the uh I want it to be abstracted away from what I'm worrying about dayto-day >> totally I could uh describe maybe a common flow that we see that's like distillation or uh general model specific model. So if you think about it um in a lot of cases it's like if the the challenge I like to pose that's like a thought experiment is like if you tell someone think of any object and it's like think of an object in your fridge the latter of those is actually easier and the latter of those is actually where a lot of models kind of get deployed into real world settings. So, for example, um you know, Monterey Bay and the Monterey Bay Aquarium Research Institute, Embari they're called. They discovered a new fish species recently and they build with RobFlow to process all of the underwater um footage that they capture from their deep sea um submarines. And they use large models like segmented anything 3 and LMS as judge to basically say hey let's take all this video footage and let's build a autolel pipeline to understand as many things as we can about all the video footage that we've collected but then ultimately you know Sam knows everything about uh you know things from fish to maybe architectural diagrams to um items in your fridge and everything in between. But they only care about things that are you know in this case underwater deep sea exploration. And so a very common flow that we see for like the fine-tune last mile distillation is okay let's take the large context of an array of models have those maybe all have a pass at understanding something use LMS as judge to say hey we agree with some amount of consensus and now we have a specialized data set then we can use that specific model and actually run that on the submarine in real time and also post-processing for faster video and if you think about it like a lot of problems are that shape where it's like yeah I want like general as much intelligence to understand the thing and then ultimately the problem I'm solving is specific enough even if not fully unbounded. And a mistake I see pretty frequently is people thinking like SAM 3 which is an awesome model and a great family of models. It's like okay well I should take SAM 3 and then maybe just fine-tune SAM 3. And in some ways that actually doesn't make as much sense because you lose the thing that makes SAM 3 awesome which is the open vocabulary capabilities. What might be better is like if you know you're distilling down to a specific fixed class list then you can actually drop the large expensive autoenccoder portion of SAM and use a specific maybe like DTOR or more specialized model depending on the task that you're solving and get all the benefits of you know speed up and accuracy uh while still having the general knowledge of preparing and curating your problem and I think a lot of problems are of that shape. >> Are you are you managing that pipeline for your customer? the yeah the tooling makes it so they can do that. Um >> do you do it on their behalf? >> Uh >> or do you give them the tooling to do it? >> Give them the platform to do that and then there's like recipes where someone can go and do that and then like any good AI company there's an FTE that if you want I can sell you get around back here and they can do it for you. That's a great example basically of the use cases and workflows that I was talking about when you're trying to find like you know collect the data as you go for your business for your enterprise and then decide how you're going to fine-tune a model a small specialized model on those use cases as you grow that's how you become more token efficient. >> Yeah. Yeah. I I I think this year and and next year you're going to see a lot of using these like monster frontier models to bootstrap, you know, like a more efficient setup that runs on open source. >> And I think, you know, that's great. Like I think this is how a lot of the open source models have been built right now. And it's it's it's proving to be like quite hard to, you know, stop that. And I would just encourage people to just you know move away potentially from like you know these uh frontier models but like you know use them use them for what for for bootstrapping that right so like um >> I love that you said that actually because that's how the word AI engineer got coined right when Swix released that blog post that coined it the way that we used to build AI products was we would start with the with the machine learning we would start by training a model then we would go try to discover a use case and what these big models allowed us to do is actually flip flip it. We said we could start with discovering a use case and then get into ML if it makes sense. And that was actually uh honestly amazing foresight from Swix because he kind of defined that pattern uh for us a few years ago and gave this amazing conference as well. >> Yeah, it's crazy how far this conference has come just three years. >> Yeah, totally. I have a question. Do you guys have questions in the in the crowd? We have a few more minutes and I know we were talking about opening this up if you guys want to ask the panelists directly. >> Just just shout. Yeah, please go ahead. >> What are the big open problems in local? I feel like you're the luminaries of the field, you know. >> You want to repeat the question? Repeat the one. You just asked what are the open problems in local AI? >> I I can repeat it. It's what are the biggest open problems in local AI? Uh >> uh it continues to be optimizations and u for inference. It continues to be getting things easily kickstarted which is you know what ODS is XOS for specific hardware. It continues to be how can I make the most out of my budget constraints and hardware constraints and how can I squeeze the most performance out of that uh you know the the models that you know I still have the same 3090s that I used to run lama 2 and like dolphin fine tunes on and now they're running coins 3.5 coins 3.6 6 27 billion parameters with excellent performance you know more on that in my next presentation. So you know uh it comes down to the optimization and we are still very early that there is space for so many players for so many contributors. We need all the help we can get to make this thing the success that it needs to be that we need to make local AI the default. This has been my stance for years now. I have been saying open source AI must win for like since forever. And the way we do that is by giving the people whether that's individuals at home, medalsized businesses or enterprises an easy way to use these models in a very efficient way that doesn't give them headaches more than solve their problems. >> Totally. And if you look at the panels that are that are happening today, those are what we believe to be the biggest open-ended questions, which is why we try to assemble this panel. And so the you know we have quantization so all about talking about the footprint so that the models do fit on these uh uh on these smaller hardware footprints. Uh there's model routing there's models generally uh so those are kind of the what we feel like are the big problems now and as we do future local AI summits um I you know every time the panels should change to be the topics dour that are kind of holding the or or going to be going to usher the next chapter in. >> Can I add one thing to that? One of the biggest challenge so I think there's two two big challenges One is what we've been talking about of basically the trade-off of simplicity versus customizability. It's local, it's yours, you can do different things, optimization, if it's hosted, it's built out of the box. The way that trade-off is is always difficult. The second, which I actually encourage people in this room to help solve is the importance of open models is becoming increasingly in question. And I actually think that like if you think local AI is important, then you think open source AI is important. And it's actually really important to be an advocate for being able to use, change, adapt, and toy with models. And so I think that that's a problem that um could increasingly be something that we feel less control over absent advocating. That's a great point. I think everyone here feels very passionate about open source. That is why we actually have access to the space at all. That's why the space is has progressed. It's a it's a necessary competitive environment. It allows for the best ideas to make their way to everybody. Um, so definitely when there's talks about that being a threat, you know, we need to invest and advocate for open source and open source is so much bigger than AI, right? Like there's a reason why computation like compute was invented in on the east coast, but Silicon Valley happened here and it's because hippies realize that they could share ideas for free in software. I think Silicon Valley is this kind of tension between the capitalists that want to make money and hippies that want to give these ideas away for free. And that tension is what creates such an incredible environment here. >> They can coexist, by the way. you can have consumers and you can sell to businesses >> and that's when this space works works best. >> Can I just say like if you care about this and you want to be more of an active participant but maybe you know you don't want to get involved in the technical side you're not technical then there are ways to get involved. So there's a website that just came out called right to intelligence.org and this is a way for you to like get involved and actually like advocate for opensource and to ensure that you know we maintain freedom of intelligence. >> Awesome. So we're out of time. I want to thank you guys so much. Locally Summit's going to be a ton of fun. Thank you guys for this incredible way to kick this off. We're gonna have amazing demos. We're running foundational intelligence models here inside of this room. Uh we're gonna have a couple of amazing uh a few more panels. It's going to be an exciting day. Um and we're all lingering here, so ask questions and please participate. Thank you. [applause] >> [music]