Close Menu
    Trending
    • Implementing IBCS rules in Power BI
    • What comes next for AI copyright lawsuits?
    • Why PDF Extraction Still Feels LikeHack
    • GenAI Will Fuel People’s Jobs, Not Replace Them. Here’s Why
    • Millions of websites to get ‘game-changing’ AI bot blocker
    • I Worked Through Labor, My Wedding and Burnout β€” For What?
    • Cloudflare will now block AI bots from crawling its clients’ websites by default
    • πŸš— Predicting Car Purchase Amounts with Neural Networks in Keras (with Code & Dataset) | by Smruti Ranjan Nayak | Jul, 2025
    AIBS News
    • Home
    • Artificial Intelligence
    • Machine Learning
    • AI Technology
    • Data Science
    • More
      • Technology
      • Business
    AIBS News
    Home»Machine Learning»ETL: Query and consolidate a large number of small JSON files with AWS Athena | by Wenjing Liu | Jan, 2025
    Machine Learning

    ETL: Query and consolidate a large number of small JSON files with AWS Athena | by Wenjing Liu | Jan, 2025

    Team_AIBS NewsBy Team_AIBS NewsJanuary 18, 2025No Comments1 Min Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Share
    Facebook Twitter LinkedIn Pinterest Email


    Wenjing Liu

    Β·

    Follow

    1 min learn

    Β·

    Simply now

    —

    πŸ‘‰ Necessities: The Amazon Bin Picture Dataset incorporates over 500,000 photos, every with a metadata JSON file sized round 1–3KB. Nevertheless, smaller information (e.g., lower than 1MB) will be inefficient to course of in bulk, as Spark incurs overhead when opening and processing every file individually. On this demo, we use AWS Athena with Trino SQL (Amazon Ion Hive SerDe) to question and consolidate 17.9MB of 10,441 JSON information into 3.9MB of 21 SNAPPY-compressed Parquet information.

    πŸ‘‰ To breed the consequence: run the next two steps: First, obtain a portion of the metadata JSON information from the general public S3 bucket to your native machine and add them to your individual S3 bucket. Then, run Trino SQL instructions in AWS Athena.

    1. ETL β€” https://github.com/nov05/udacity-nd009t-capstone-starter/blob/master/starter/ETL.ipynb

    2. AWS Athena Trino SQL β€” https://github.com/nov05/udacity-nd009t-capstone-starter/blob/master/starter/AWS%20Athena%20Trino%20SQL.md



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleInstagram and YouTube Prepare to Benefit From a TikTok Ban
    Next Article Learning from Machine Learning | Sebastian Raschka: Mastering ML and Pushing AI Forward Responsibly | by Seth Levine
    Team_AIBS News
    • Website

    Related Posts

    Machine Learning

    Why PDF Extraction Still Feels LikeHack

    July 1, 2025
    Machine Learning

    πŸš— Predicting Car Purchase Amounts with Neural Networks in Keras (with Code & Dataset) | by Smruti Ranjan Nayak | Jul, 2025

    July 1, 2025
    Machine Learning

    Reinforcement Learning in the Age of Modern AI | by @pramodchandrayan | Jul, 2025

    July 1, 2025
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Implementing IBCS rules in Power BI

    July 1, 2025

    I Tried Buying a Car Through Amazon: Here Are the Pros, Cons

    December 10, 2024

    Amazon and eBay to pay ‘fair share’ for e-waste recycling

    December 10, 2024

    Artificial Intelligence Concerns & Predictions For 2025

    December 10, 2024

    Barbara Corcoran: Entrepreneurs Must ‘Embrace Change’

    December 10, 2024
    Categories
    • AI Technology
    • Artificial Intelligence
    • Business
    • Data Science
    • Machine Learning
    • Technology
    Most Popular

    Why Oil, Gas Giant Chevron Is Laying Off Up to 8,000 Workers

    February 13, 2025

    Hegseth orders pause in offensive US cyber operations against Russia

    March 3, 2025

    Meet a Swift Student Challenge Winner Attending Apple’s WWDC

    June 10, 2025
    Our Picks

    Implementing IBCS rules in Power BI

    July 1, 2025

    What comes next for AI copyright lawsuits?

    July 1, 2025

    Why PDF Extraction Still Feels LikeHack

    July 1, 2025
    Categories
    • AI Technology
    • Artificial Intelligence
    • Business
    • Data Science
    • Machine Learning
    • Technology
    • Privacy Policy
    • Disclaimer
    • Terms and Conditions
    • About us
    • Contact us
    Copyright Β© 2024 Aibsnews.comAll Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.