I joined Lemmy back in 2020 and have been using it as @qaz@lemmy.ml until somewhere in 2023 when I switched to lemmy.world. I’m interested in systemd/Linux, FOSS, and Selfhosting.

  • 104 Posts
  • 531 Comments
Joined vor 3 Jahren
cake
Cake day: 10. Juni 2023

help-circle













  • This reminds me of something I sometimes see in shows on like Netflix and other media. I can’t remember a specific example, but you often have generic anti-capitalist comments from characters (often portrayed as edgy). It often feels a bit, artificial, like a “fellow kids” moment but politically, I guess? Like activism as a prop / character trait, inserted into a multi-million media production. Maybe someone else can better put it to words, if I had more time I would’ve written a shorter comment


  • Deepseek recently published a paper in which they describe that vision tokens contain more information than text tokens and that this can be used to compress context.

    We present DeepSeek-OCR as an initial investigation into the feasibility of compressing long contexts via optical 2D mapping.

    Experiments show that when the number of text tokens is within 10 times that of vision tokens (i.e., a compression ratio < 10×), the model can achieve decoding (OCR) precision of 97%. Even at a compression ratio of 20×, the OCR accuracy still remains at about 60%. This shows considerable promise for research areas such as historical long-context compression and memory forgetting mechanisms in LLMs.

    It reminds me of LLM caveman speak, it used to have another option to use Chinese instead of English. A language like Chinese is seemingly better at encoding information in fewer tokens and I think this is the same mechanism why OCR tokens work so well.

    That said, I also doubt that voice messages are more efficient than text prompts, but it’s best not to waste too much time engaging with these sorts of LinkedIn posts (and LinkedIn in general).








  • Some communities don’t have specific rules, and some have several. It can be quite hard to judge why something is breaking rules. On the video’s community for example people often report videos, so I have to go through the entire thing to figure out what’s wrong with it (and some people post videos that are more than an hour). Sometimes content is reported because of something the people who created it did, something that is not apparent when watching the video.



  • I’d say go ahead but make sure it produces accurate enough results and make sure to add something like [AI Transcribed] in front so people can take the potential for additional errors into consideration when reading it.

    Also, if you’re using an online service make sure you’re using something that doesn’t use it as training data. Many (probably almost all) artists / photographers won’t appreciate that.