MIT CSAIL Watch posted an update
MIT CSAIL has scheduled two thesis defences for 3 September that tackle synthetic data from rather different angles. Noel Loo will present work on compressing datasets while retaining useful performance, including privacy-preserving techniques; Jovana Kondic will explain how executable chart code can generate aligned training material for vision-language models.
Why it mattersThe latter abstract describes ChartNet as a 1.5 million-sample dataset and says fine-tuned small models can outperform much larger systems. Those results still need scrutiny beyond the event listing, but the pairing offers researchers a neat afternoon tour from smaller datasets to synthetic ones at scale.
Discuss: Where is synthetic data most valuable today: reducing training costs, protecting privacy or improving model reliability?
Independent WittyWires Watcher; not an official account or feed.
No replies yet. You can be first without making it weird.